Emotion Recognition System From Speech And Visual Information Based On Convolutional Neural Networks
2020 Β· Nicolae-Catalin Ristea, Liviu Cristian Dutu, Anamaria Radoi
Abstract
Emotion recognition has become an important field of research in the human-computer interactions domain. The latest advancements in the field show that combining visual with audio information lead to better results if compared to the case of using a single source of information separately. From a visual point of view, a human emotion can be recognized by analyzing the facial expression of the person. More precisely, the human emotion can be described through a combination of several Facial Action Units. In this paper, we propose a system that is able to recognize emotions with a high accuracy rate and in real time, based on deep Convolutional Neural Networks. In order to increase the accuracy of the recognition system, we analyze also the speech data and fuse the information coming from both sources, i.e., visual and audio. Experimental results show the effectiveness of the proposed scheme for emotion recognition and the importance of combining visual with audio data.
Authors
(none)
Tags
Stats
Related papers
- Multimodal Fusion With Deep Neural Networks For Audio-video Emotion Recognition (2019)0.00
- Attention Based Fully Convolutional Network For Speech Emotion Recognition (2018)15.25
- Emotion Recognition From Speech (2019)0.00
- Emodiarize: Speaker Diarization And Emotion Identification From Speech Signals Using Convolutional Neural Networks (2023)0.00
- Emotech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information With Hybrid Recurrent Network (2025)8.35
- Temporal Aggregation Of Audio-visual Modalities For Emotion Recognition (2020)8.09
- Interpretable Multimodal Emotion Recognition Using Hybrid Fusion Of Speech And Image Data (2022)11.85
- Audio Visual Emotion Recognition With Temporal Alignment And Perception Attention (2016)0.00