Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Emotion Recognition with Intensity Level from Bangla Speech using Feature Transformation and Cascaded Deep Learning Model

Author 1: Md. Masum Billah Author 2: Md. Likhon Sarker Author 3: M. A. H. Akhand Author 4: Md Abdus Samad Kamal
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 15, No. 4 · Published 2024 · Cited by 11

DOI: https://doi.org/10.14569/IJACSA.2024.0150460

Abstract

Speech Emotion Recognition (SER) identifies and categorizes emotional states by analyzing speech signals. The intensity of specific emotional expressions (e.g., anger) conveys critical directives and plays a crucial role in social behavior. SER is intrinsically language-specific; this study investigated a novel cascaded deep learning (DL) model to Bangla SER with intensity level. The proposed method employs the Mel-Frequency Cepstral Coefficient, Short-Time Fourier Transform (STFT), and Chroma STFT signal transformation techniques; the respective trans-formed features are blended into a 3D form and used as the input of the DL model. The cascaded model performs the task in two stages: classify emotion in Stage 1 and then measure the intensity in Stage 2. DL architecture used in both stages is the same, which consists of a 3D Convolutional Neural Network (CNN), a Time Distribution Flatten (TDF) layer, a Long Short-term Memory (LSTM), and a Bidirectional LSTM (Bi-LSTM). CNN first extracts features from 3D formed input; the features are passed through the TDF layer, Bi-LSTM, and LSTM; finally, the model classifies emotion along with its intensity level. The proposed model has been evaluated rigorously using developed KBES and other datasets. The proposed model revealed as the best-suited SER method compared to existing prominent methods achieving accuracy of 88.30% and 71.67% for RAVDESS and KBES datasets, respectively.

Keywords

How to Cite this Article

Billah, M. M., Sarker, M. L., Akhand, M. A. H., & Kamal, M. A. S. (2024). Emotion Recognition with Intensity Level from Bangla Speech using Feature Transformation and Cascaded Deep Learning Model. International Journal of Advanced Computer Science and Applications, 15(4). https://doi.org/10.14569/IJACSA.2024.0150460

Billah, Md. Masum, et al.. "Emotion Recognition with Intensity Level from Bangla Speech using Feature Transformation and Cascaded Deep Learning Model." International Journal of Advanced Computer Science and Applications, vol. 15, no. 4, 2024, https://doi.org/10.14569/IJACSA.2024.0150460.

@article{Billah2024,
  title     = {Emotion Recognition with Intensity Level from Bangla Speech using Feature Transformation and Cascaded Deep Learning Model},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {15},
  number    = {4},
  year      = {2024},
  publisher = {The Science and Information Organization},
  author    = {Md. Masum Billah and Md. Likhon Sarker and M. A. H. Akhand and Md Abdus Samad Kamal},
  doi       = {10.14569/IJACSA.2024.0150460},
  url       = {https://doi.org/10.14569/IJACSA.2024.0150460}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.