MSP-Podcast
Emerging16papers using it
2021first seen
The MSP-Podcast dataset is a benchmark that contains in-the-wild speech data used to evaluate the effectiveness of speech emotion conversion methods.
Papers using MSP-Podcast (16)
- Multistage Linguistic Conditioning Of Convolutional Layers For Speech Emotion RecognitionTargetSEC: Plug-and-Play In-the-Wild Speech Emotion Conversion via Arousal-Conditioned Latent Style DiffusionAdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge IntegrationJoint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion RecognitionEnhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration ModelingSpeech Emotion Recognition with ASR Transcripts: A Comprehensive Study
on Word Error Rate and Fusion TechniquesMultistage linguistic conditioning of convolutional layers for speech
emotion recognitionMSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal
Self-Supervised Learning for Speech Emotion RecognitionSentiment-Aware Automatic Speech Recognition pre-training for enhanced
Speech Emotion RecognitionIn-the-wild Speech Emotion Conversion Using Disentangled Self-Supervised
Representations and Neural Vocoder-based ResynthesisPersonalized Adaptation with Pre-trained Speech Encoders for Continuous
Emotion RecognitionRepresentation learning through cross-modal conditional teacher-student
training for speech emotion recognitionNon-Contrastive Self-Supervised Learning of Utterance-Level Speech
RepresentationsDawn of the transformer era in speech emotion recognition: closing the
valence gapemoDARTS: Joint Optimisation of CNN & Sequential Neural Network
Architectures for Superior Speech Emotion RecognitionFusion approaches for emotion recognition from speech using acoustic and
text-based features