Capturing Long-term Temporal Dependencies With Convolutional Networks For Continuous Emotion Recognition
2017 Β· Soheil Khorram, Zakaria Aldeneh, Dimitrios Dimitriadis, et al.
Abstract
The goal of continuous emotion recognition is to assign an emotion value to every frame in a sequence of acoustic features. We show that incorporating long-term temporal dependencies is critical for continuous emotion recognition tasks. To this end, we first investigate architectures that use dilated convolutions. We show that even though such architectures outperform previously reported systems, the output signals produced from such architectures undergo erratic changes between consecutive time steps. This is inconsistent with the slow moving ground-truth emotion labels that are obtained from human annotators. To deal with this problem, we model a downsampled version of the input signal and then generate the output signal through upsampling. Not only does the resulting downsampling/upsampling network achieve good performance, it also generates smooth output trajectories. Our method yields the best known audio-only performance on the RECOLA dataset.
Authors
(none)
Tags
Stats
Related papers
- Jointly Aligning And Predicting Continuous Emotion Annotations (2019)8.60
- Dynamic Time-alignment Of Dimensional Annotations Of Emotion Using Recurrent Neural Networks (2022)0.00
- Attention Based Fully Convolutional Network For Speech Emotion Recognition (2018)15.25
- Unifying The Discrete And Continuous Emotion Labels For Speech Emotion Recognition (2022)0.00
- Continuous Multimodal Emotion Recognition Approach For AVEC 2017 (2017)0.00
- Recursive Joint Attention For Audio-visual Fusion In Regression Based Emotion Recognition (2023)9.59
- Multi-time-scale Convolution For Emotion Recognition From Speech Audio Signals (2020)11.67
- Multi-modal Continuous Valence And Arousal Prediction In The Wild Using Deep 3D Features And Sequence Modeling (2020)0.00