Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN

P V V S Srinivas; Pragnyaban Mishra

doi:10.14569/IJACSA.2022.0130172

DOI: 10.14569/IJACSA.2022.0130172

PDF

Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN

Author 1: P V V S Srinivas

Author 2: Pragnyaban Mishra

International Journal of Advanced Computer Science and Applications(IJACSA), Volume 13 Issue 1, 2022.

Abstract and Keywords
How to Cite this Article
{} BibTeX Source

Abstract: This Emotion recognition plays a prominent role in today's intelligent system applications. Human computer interface, health care, law, and entertainment are a few of the applications where emotion recognition is used. Humans convey their emotions in the form of text, voice, and facial expressions, thus developing a multimodal emotional recognition system playing a crucial role in human-computer or intelligent system communication. The majority of established emotional recognition algorithms only identify emotions in unique data, such as text, audio, or image data. A multimodal system uses information from a variety of sources and fuses the information by using fusion techniques and categories to improve recognition accuracy. In this paper, a multimodal system to recognise emotions was presented that fuses the features from information obtained from heterogenous modalities like audio and video. For audio feature extraction energy, zero crossing rate and Mel-Frequency Cepstral Coefficients (MFCC) techniques are considered. Of these, MFCC produced promising results. For video feature extraction, first the videos are converted to frames and stored in a linear scale space by using a spatial temporal Gaussian Kernel. The features from the images are further extracted by applying a Gaussian weighted function to the second momentum matrix of linear scale space data. The Marginal Fisher Analysis (MFA) fusion method is used to fuse both the audio and video features, and the resulted features are given to the FERCNN model for evaluation. For experimentation, the RAVDESS and CREMAD datasets, which contain audio and video data, are used. Accuracy levels of 95.56, 96.28, and 95.07 on the RAVDESS dataset and accuracies of 80.50, 97.88, and 69.66 on the CREMAD dataset in audio, video, and multimodal modalities are achieved, whose performance is better than the existing multimodal systems.

Keywords: Emotion recognition; multimodal; fusion; MFCC; MFA; FERCNN; CREMAD; RAVDESS

P V V S Srinivas and Pragnyaban Mishra, “Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN” International Journal of Advanced Computer Science and Applications(IJACSA), 13(1), 2022. http://dx.doi.org/10.14569/IJACSA.2022.0130172

@article{Srinivas2022,
title = {Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN},
journal = {International Journal of Advanced Computer Science and Applications},
doi = {10.14569/IJACSA.2022.0130172},
url = {http://dx.doi.org/10.14569/IJACSA.2022.0130172},
year = {2022},
publisher = {The Science and Information Organization},
volume = {13},
number = {1},
author = {P V V S Srinivas and Pragnyaban Mishra}
}

Copyright Statement: This is an open access article licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, even commercially as long as the original work is properly cited.

Human Emotion Recognition by Integrating Facial and Speech Features: An Implementation of Multimodal Framework using CNN

Upcoming Conferences