Time-frequency Transformer: A Novel Time Frequency Joint Learning Method For Speech Emotion Recognition
2023 Β· Yong Wang, Cheng Lu, Yuan Zong, et al.
Abstract
In this paper, we propose a novel time-frequency joint learning method for speech emotion recognition, called Time-Frequency Transformer. Its advantage is that the Time-Frequency Transformer can excavate global emotion patterns in the time-frequency domain of speech signal while modeling the local emotional correlations in the time domain and frequency domain respectively. For the purpose, we first design a Time Transformer and Frequency Transformer to capture the local emotion patterns between frames and inside frequency bands respectively, so as to ensure the integrity of the emotion information modeling in both time and frequency domains. Then, a Time-Frequency Transformer is proposed to mine the time-frequency emotional correlations through the local time-domain and frequency-domain emotion features for learning more discriminative global speech emotion representation. The whole process is a time-frequency joint learning process implemented by a series of Transformer models. Experi
Authors
(none)
Tags
Stats
Related papers
- Learning Local To Global Feature Aggregation For Speech Emotion Recognition (2023)8.09
- Emoformer: A Text-independent Speech Emotion Recognition Using A Hybrid Transformer-cnn Model (2025)6.34
- Multi-modal Emotion Recognition By Text, Speech And Video Using Pretrained Transformers (2024)0.00
- Key-sparse Transformer For Multimodal Speech Emotion Recognition (2021)13.50
- Tf-locoformer: Transformer With Local Modeling By Convolution For Speech Separation And Enhancement (2024)10.35
- Speech Emotion Recognition Via Cnn-transformer And Multidimensional Attention Mechanism (2024)0.00
- A Transfer Learning Method For Speech Emotion Recognition From Automatic Speech Recognition (2020)0.00
- Speech Swin-transformer: Exploring A Hierarchical Transformer With Shifted Windows For Speech Emotion Recognition (2024)11.29