Facebook pixel tracking

The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |

Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN

Author 1: P V V S Srinivas Author 2: Pragnyaban Mishra
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 13, No. 1 · Published 2022 · Cited by 10

DOI: https://doi.org/10.14569/IJACSA.2022.0130172

Abstract

This Emotion recognition plays a prominent role in today's intelligent system applications. Human computer interface, health care, law, and entertainment are a few of the applications where emotion recognition is used. Humans convey their emotions in the form of text, voice, and facial expressions, thus developing a multimodal emotional recognition system playing a crucial role in human-computer or intelligent system communication. The majority of established emotional recognition algorithms only identify emotions in unique data, such as text, audio, or image data. A multimodal system uses information from a variety of sources and fuses the information by using fusion techniques and categories to improve recognition accuracy. In this paper, a multimodal system to recognise emotions was presented that fuses the features from information obtained from heterogenous modalities like audio and video. For audio feature extraction energy, zero crossing rate and Mel-Frequency Cepstral Coefficients (MFCC) techniques are considered. Of these, MFCC produced promising results. For video feature extraction, first the videos are converted to frames and stored in a linear scale space by using a spatial temporal Gaussian Kernel. The features from the images are further extracted by applying a Gaussian weighted function to the second momentum matrix of linear scale space data. The Marginal Fisher Analysis (MFA) fusion method is used to fuse both the audio and video features, and the resulted features are given to the FERCNN model for evaluation. For experimentation, the RAVDESS and CREMAD datasets, which contain audio and video data, are used. Accuracy levels of 95.56, 96.28, and 95.07 on the RAVDESS dataset and accuracies of 80.50, 97.88, and 69.66 on the CREMAD dataset in audio, video, and multimodal modalities are achieved, whose performance is better than the existing multimodal systems.

Keywords

How to Cite this Article

Srinivas, P. V. V. S., & Mishra, P. (2022). Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN. International Journal of Advanced Computer Science and Applications, 13(1). https://doi.org/10.14569/IJACSA.2022.0130172

Srinivas, P V V S, and Pragnyaban Mishra. "Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN." International Journal of Advanced Computer Science and Applications, vol. 13, no. 1, 2022, https://doi.org/10.14569/IJACSA.2022.0130172.

@article{Srinivas2022,
  title     = {Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {13},
  number    = {1},
  year      = {2022},
  publisher = {The Science and Information Organization},
  author    = {P V V S Srinivas and Pragnyaban Mishra},
  doi       = {10.14569/IJACSA.2022.0130172},
  url       = {https://doi.org/10.14569/IJACSA.2022.0130172}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.