MELD
Emerging18papers using it
2022first seen
MELD is a dataset that contains multi-modal emotional dialogues used to evaluate the emotional expressiveness and prosodic features in speech synthesis.
Papers using MELD (18)
- A Mixture-of-Experts Model for Multimodal Emotion Recognition in ConversationsRLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech SynthesisOn the Contribution of Lexical Features to Speech Emotion RecognitionEmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language ModelTowards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion RecognitionVesper: A Compact and Effective Pretrained Model for Speech Emotion
RecognitionExtending RNN-T-based speech recognition systems with emotion and
language classificationDST: Deformable Speech Transformer for Emotion Recognitiondeep learning of segment-level feature representation for speech emotion
recognition in conversationsDWFormer: Dynamic Window transFormer for Speech Emotion RecognitionA Change of Heart: Improving Speech Emotion Recognition through
Speech-to-Text Modality ConversionSpeechFormer: A Hierarchical Efficient Framework Incorporating the
Characteristics of SpeechASR and Emotional Speech: A Word-Level Investigation of the Mutual
Impact of Speech and Emotion RecognitionMELD-ST: An Emotion-aware Speech Translation DatasetMulti-Scale Temporal Transformer For Speech Emotion RecognitionWavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion RecognitionTemporal-Frequency State Space Duality: An Efficient Paradigm for Speech
Emotion RecognitionSpeechFormer++: A Hierarchical Efficient Framework for Paralinguistic
Speech Processing