From Weak Labels To Strong Results: Utilizing 5,000 Hours Of Noisy Classroom Transcripts With Minimal Accurate Data
2025 Β· Ahmed Adel Attia, Dorottya Demszky, Jing Liu, et al.
Abstract
Recent progress in speech recognition has relied on models trained on vast amounts of labeled data. However, classroom Automatic Speech Recognition (ASR) faces the real-world challenge of abundant weak transcripts paired with only a small amount of accurate, gold-standard data. In such low-resource settings, high transcription costs make re-transcription impractical. To address this, we ask: what is the best approach when abundant inexpensive weak transcripts coexist with limited gold-standard data, as is the case for classroom speech data? We propose Weakly Supervised Pretraining (WSP), a two-step process where models are first pretrained on weak transcripts in a supervised manner, and then fine-tuned on accurate data. Our results, based on both synthetic and real weak transcripts, show that WSP outperforms alternative methods, establishing it as an effective training methodology for low-resource ASR in real-world scenarios.
Authors
(none)
Tags
Stats
Related papers
- Weakly-supervised Speech Pre-training: A Case Study On Target Speech Recognition (2023)8.09
- Large Scale Weakly And Semi-supervised Learning For Low-resource Video ASR (2020)0.00
- Training ASR Models By Generation Of Contextual Information (2019)0.00
- Leveraging Weakly Supervised Data To Improve End-to-end Speech-to-text Translation (2018)13.05
- Learning From Flawed Data: Weakly Supervised Automatic Speech Recognition (2023)13.45
- Whistle: Data-efficient Multilingual And Crosslingual Speech Recognition Via Weakly Phonetic Supervision (2024)10.38
- Speechnet: Weakly Supervised, End-to-end Speech Recognition At Industrial Scale (2022)0.00
- Pretraining By Backtranslation For End-to-end ASR In Low-resource Settings (2018)0.00