YouTube
Emerging7papers using it
2022first seen
The 'YouTube' dataset contains audio recordings of spoken language and is used to evaluate and improve the performance of automatic speech recognition systems, particularly in recognizing different varieties of English, including African American English.
Papers using YouTube (7)
- A Vector Quantized Approach for Text to Speech Synthesis on Real-World
Spontaneous SpeechEfficient Domain Adaptation for Speech Foundation ModelsSpoken Language Identification System for English-Mandarin
Code-Switching Child-Directed SpeechE2E Segmenter: Joint Segmenting and Decoding for Long-Form ASREnd-to-End Multi-Person Audio/Visual Automatic Speech RecognitionExperiments on Turkish ASR with Self-Supervised Speech Representation
LearningImproving Speech Recognition for African American English With Audio
Classification