IEMOCAP
Emerging75papers using it
2021first seen
IEMOCAP is a dataset used for evaluating speech emotion recognition systems, containing labeled speech data that includes various emotional expressions.
Papers using IEMOCAP (75)
- Emotion2vec: Self-supervised Pre-training For Speech Emotion RepresentationMultistage Linguistic Conditioning Of Convolutional Layers For Speech Emotion RecognitionEmoHRNet: High-Resolution Neural Network Based Speech Emotion RecognitionLeveraging Cross-Attention Transformer and Multi-Feature Fusion for
Cross-Linguistic Speech Emotion RecognitionAligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement LearningAttention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion RecognitionMulti-Channel Speech Enhancement for Cocktail Party Speech Emotion RecognitionA Mixture-of-Experts Model for Multimodal Emotion Recognition in ConversationsMulti-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level AttentionEnhancing Speech Emotion Recognition via Fine-Tuning Pre-Trained Models and Hyper-Parameter OptimisationEmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language ModelEmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion RecognitionSpeech Emotion Recognition via Entropy-Aware Score SelectionEnhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker DiarizationTowards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion RecognitionQieemo: Speech Is All You Need in the Emotion Recognition in
ConversationsMetadata-Enhanced Speech Emotion Recognition: Augmented Residual
Integration and Co-Attention in Two-Stage Fine-TuningSpeech Emotion Recognition with ASR Transcripts: A Comprehensive Study
on Word Error Rate and Fusion TechniquesA Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion
Recognition, Speaker Verification and Spoken Language UnderstandingMingling or Misalignment? Temporal Shift for Speech Emotion Recognition
with Pre-trained RepresentationsExploring Wav2vec 2.0 fine-tuning for improved speech emotion
recognitionMFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and
Hierarchical Cooperative AttentionLayer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study
on Speech Emotion RecognitionThe Role of Phonetic Units in Speech Emotion RecognitionUsing Large Pre-Trained Models with Cross-Modal Attention for
Multi-Modal Emotion RecognitionMultistage linguistic conditioning of convolutional layers for speech
emotion recognitionSpeech Emotion Recognition Via CNN-Transformer and Multidimensional
Attention MechanismActive Learning with Task Adaptation Pre-training for Speech Emotion
RecognitionFusing ASR Outputs in Joint Training for Speech Emotion RecognitionSpeech Emotion Recognition using Self-Supervised FeaturesSpeechEQ: Speech Emotion Recognition based on Multi-scale Unified
Datasets and Multitask LearningVesper: A Compact and Effective Pretrained Model for Speech Emotion
RecognitionLight-SERNet: A lightweight fully convolutional neural network for
speech emotion recognitionMMER: Multimodal Multi-task Learning for Speech Emotion RecognitionCTA-RNN: Channel and Temporal-wise Attention RNN Leveraging Pre-trained
ASR Embeddings for Speech Emotion RecognitionExtending RNN-T-based speech recognition systems with emotion and
language classificationDST: Deformable Speech Transformer for Emotion RecognitionReal-time Speech Emotion Recognition Based on Syllable-Level Feature
ExtractionDWFormer: Dynamic Window transFormer for Speech Emotion RecognitionLeveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion
RecognitionFrame-level emotional state alignment method for speech emotion
recognitionExploring Multilingual Unseen Speaker Emotion Recognition: Leveraging
Co-Attention Cues in Multitask LearningPCQ: Emotion Recognition in Speech via Progressive Channel QueryingA Cross-Corpus Speech Emotion Recognition Method Based on Supervised
Contrastive LearningRepresentation learning through cross-modal conditional teacher-student
training for speech emotion recognitionSpeaker Normalization for Self-supervised Speech Emotion RecognitionTowards a Common Speech Analysis EngineSpeechFormer: A Hierarchical Efficient Framework Incorporating the
Characteristics of SpeechSemi-FedSER: Semi-supervised Learning for Speech Emotion Recognition On
Federated Learning using Multiview Pseudo-LabelingSpeech Emotion Recognition with Co-Attention based Multi-level Acoustic
InformationNeural Architecture Search for Speech Emotion RecognitionMultitask Learning from Augmented Auxiliary Data for Improving Speech
Emotion RecognitionNon-Contrastive Self-Supervised Learning of Utterance-Level Speech
RepresentationsOn the Efficacy and Noise-Robustness of Jointly Learned Speech Emotion
and Automatic Speech RecognitionEnhancing Speech Emotion Recognition Through Differentiable Architecture
SearchASR and Emotional Speech: A Word-Level Investigation of the Mutual
Impact of Speech and Emotion RecognitionGEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio
Pretraining for Accurate Speech Emotion RecognitionTime-Frequency Transformer: A Novel Time Frequency Joint Learning Method
for Speech Emotion RecognitionIntegrating Contrastive Learning into a Multitask Transformer Model for
Effective Domain AdaptationMulti-Level Knowledge Distillation for Speech Emotion Recognition in
Noisy ConditionsDeep functional multiple index models with an application to SERGMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer
Learning for Speech Emotion RecognitionSELM: Enhancing Speech Emotion Recognition for Out-of-Domain ScenariosMulti-Scale Temporal Transformer For Speech Emotion RecognitionEnhancing Speech Emotion Recognition through Segmental Average Pooling
of Self-Supervised Learning FeaturesEnd-to-End Integration of Speech Emotion Recognition with Voice Activity
Detection using Self-Supervised Learning FeaturesWavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion RecognitionTemporal-Frequency State Space Duality: An Efficient Paradigm for Speech
Emotion RecognitionDawn of the transformer era in speech emotion recognition: closing the
valence gapA Graph Isomorphism Network with Weighted Multiple Aggregators for
Speech Emotion RecognitionMultimodal Speech Emotion Recognition using Cross Attention with Aligned
Audio and TextSpeechFormer++: A Hierarchical Efficient Framework for Paralinguistic
Speech ProcessingIntegrating Emotion Recognition with Speech Recognition and Speaker
Diarisation for ConversationsemoDARTS: Joint Optimisation of CNN & Sequential Neural Network
Architectures for Superior Speech Emotion RecognitionFusion approaches for emotion recognition from speech using acoustic and
text-based features