Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Speech Emotion Recognition from Audio Data Using LSTM Model

Author 1: Md. Mahbub-Or-Rashid Author 2: Akash Kumar Nondi Author 3: Abdullah Al Sadnun Author 4: Md. Anwar Hussen Wadud Author 5: T M Amir Ul Haque Bhuiyan Author 6: Md. Saddam Hossain
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 16, No. 7 · Published 2025

DOI: https://doi.org/10.14569/IJACSA.2025.0160726

Abstract

The capacity to comprehend and interact with others through language is the most valuable human ability. Since emotions are crucial to communication, we are well-trained to recognize and interpret the many emotions we encounter. Contrary to popular assumption, the subjective aspect of human mood makes emotion recognition difficult for computers. There are some works based on Emotion recognition using images, text, and audio. We are here working on the audio dataset to find the accurate human emotion for computers to understand. In this work, we have utilized a Long Short-Term Memory (LSTM) model to implement Speech Emotion Recognition (SER) from Audio data on two different datasets: the Toronto Emotional Speech Set (TESS) and the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). The accuracy rates of our LSTM-based model were impressive, with 91.25% for the RAVDESS dataset and 98.05% for the TESS dataset; the combined accuracy for both datasets was 87.66%. These results highlight the effectiveness of the LSTM model in effectively identifying and categorizing emotional states from audio files. The study adds significant knowledge to the field of speech emotion recognition by emphasizing the model’s ability to handle a variety of datasets and its potential.

Keywords

How to Cite this Article

Mahbub-Or-Rashid, M., Nondi, A. K., Sadnun, A. A., Wadud, M. A. H., Bhuiyan, T. M. A. U. H., & Hossain, M. S. (2025). Speech Emotion Recognition from Audio Data Using LSTM Model. International Journal of Advanced Computer Science and Applications, 16(7). https://doi.org/10.14569/IJACSA.2025.0160726

Mahbub-Or-Rashid, Md., et al.. "Speech Emotion Recognition from Audio Data Using LSTM Model." International Journal of Advanced Computer Science and Applications, vol. 16, no. 7, 2025, https://doi.org/10.14569/IJACSA.2025.0160726.

@article{Mahbub-Or-Rashid2025,
  title     = {Speech Emotion Recognition from Audio Data Using LSTM Model},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {16},
  number    = {7},
  year      = {2025},
  publisher = {The Science and Information Organization},
  author    = {Md. Mahbub-Or-Rashid and Akash Kumar Nondi and Abdullah Al Sadnun and Md. Anwar Hussen Wadud and T M Amir Ul Haque Bhuiyan and Md. Saddam Hossain},
  doi       = {10.14569/IJACSA.2025.0160726},
  url       = {https://doi.org/10.14569/IJACSA.2025.0160726}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.