Speech Emotion Recognition Via Cnn-transformer And Multidimensional Attention Mechanism
2024 Β· Xiaoyu Tang, Yixin Lin, Ting Dang, et al.
Abstract
Speech Emotion Recognition (SER) is crucial in human-machine interactions. Mainstream approaches utilize Convolutional Neural Networks or Recurrent Neural Networks to learn local energy feature representations of speech segments from speech information, but struggle with capturing global information such as the duration of energy in speech. Some use Transformers to capture global information, but there is room for improvement in terms of parameter count and performance. Furthermore, existing attention mechanisms focus on spatial or channel dimensions, hindering learning of important temporal information in speech. In this paper, to model local and global information at different levels of granularity in speech and capture temporal, spatial and channel dependencies in speech signals, we propose a Speech Emotion Recognition network based on CNN-Transformer and multi-dimensional attention mechanisms. Specifically, a stack of CNN blocks is dedicated to capturing local information in speech
Authors
(none)
Tags
Stats
Related papers
- Speech Emotion Recognition With Multiscale Area Attention And Data Augmentation (2021)13.65
- Searching For Effective Preprocessing Method And Cnn-based Architecture With Efficient Channel Attention On Speech Emotion Recognition (2024)2.26
- Enhanced Speech Emotion Recognition With Efficient Channel Attention Guided Deep Cnn-bilstm Framework (2024)0.00
- Cross-language Speech Emotion Recognition Using Multimodal Dual Attention Transformers (2023)0.00
- Emoformer: A Text-independent Speech Emotion Recognition Using A Hybrid Transformer-cnn Model (2025)6.34
- Attention Based Fully Convolutional Network For Speech Emotion Recognition (2018)15.25
- Speech Emotion Recognition With Co-attention Based Multi-level Acoustic Information (2022)16.17
- Leveraging Cross-attention Transformer And Multi-feature Fusion For Cross-linguistic Speech Emotion Recognition (2025)4.52