Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

A Hybrid SMOTE-Ensemble Framework for Robust Speech Emotion Recognition on Imbalanced Data

Author 1: Eyad Megdadi Author 2: Mariam Aljouhi Author 3: Zainab Alblooshi Author 4: Khaled Shaalan
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 17, No. 7 · Published 2026

DOI: https://doi.org/10.14569/IJACSA.2026.0170720

Abstract

Speech Emotion Recognition (SER) has attracted considerable attention in recent years because of its potential applications in healthcare, intelligent assistants, customer service, and human-computer interaction. However, the performance of SER systems is often affected by the imbalance of emotion classes in benchmark datasets, where minority emotions are underrepresented and difficult to classify accurately. This study reproduces the work of Sahu et al. on multimodal Speech Emotion Recognition using the Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset. Further, it proposes a Hybrid SMOTE-Ensemble Framework to improve classification performance on imbalanced data. The reproduction stage follows the original methodology by employing handcrafted acoustic features together with TF-IDF textual features under Audio-only, Text-only, and Audio+Text modalities. The enhancement stage introduces two XGBoost classifiers trained independently on the original imbalanced dataset and on a SMOTE-balanced dataset. Their prediction probabilities are combined using three ensemble strategies, namely Simple Averaging, Confidence-Based Selection, and Adaptive Weighting. In addition, the effectiveness of the proposed strategy is further investigated using a LinearSVC classifier. Experimental results demonstrate that the reproduced models achieve results that are highly consistent with those reported in the original study. Furthermore, the proposed Hybrid SMOTE-Ensemble Framework improves the recognition of minority emotion classes while maintaining balanced overall performance. The XGBoost model trained on SMOTE-balanced data achieves higher F1-score and recall than the baseline model, whereas the ensemble approaches provide additional improvements in classification performance. The LinearSVC experiments further confirm that combining oversampling and ensemble learning offers a robust solution to the class imbalance problem in Speech Emotion Recognition.

Keywords

How to Cite this Article

Megdadi, E., Aljouhi, M., Alblooshi, Z., & Shaalan, K. (2026). A Hybrid SMOTE-Ensemble Framework for Robust Speech Emotion Recognition on Imbalanced Data. International Journal of Advanced Computer Science and Applications, 17(7). https://doi.org/10.14569/IJACSA.2026.0170720

Megdadi, Eyad, et al.. "A Hybrid SMOTE-Ensemble Framework for Robust Speech Emotion Recognition on Imbalanced Data." International Journal of Advanced Computer Science and Applications, vol. 17, no. 7, 2026, https://doi.org/10.14569/IJACSA.2026.0170720.

@article{Megdadi2026,
  title     = {A Hybrid SMOTE-Ensemble Framework for Robust Speech Emotion Recognition on Imbalanced Data},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {17},
  number    = {7},
  year      = {2026},
  publisher = {The Science and Information Organization},
  author    = {Eyad Megdadi and Mariam Aljouhi and Zainab Alblooshi and Khaled Shaalan},
  doi       = {10.14569/IJACSA.2026.0170720},
  url       = {https://doi.org/10.14569/IJACSA.2026.0170720}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.