MUSAN
Emerging3papers using it
2022first seen
MUSAN: A Music, Speech, and Noise Corpus MUSAN is a corpus of music, speech, and noise recordings designed for training models for voice activity detection and music/speech discrimination. This is a comprehensive collection suitable for various audio processing tasks. Dataset Structure The dataset is organized into thr
Papers using MUSAN (3)
- Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and DecoderJoint Speaker Encoder and Neural Back-end Model for Fully End-to-End
Automatic Speaker Verification with Multiple Enrollment UtterancesOn the Efficacy and Noise-Robustness of Jointly Learned Speech Emotion
and Automatic Speech Recognition