Common Voice
Canonical48papers using it
2021first seen
Mozilla's massively-multilingual, crowd-sourced read-speech corpus for speech recognition.
Papers using Common Voice (48)
- Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech RecognitionSN-WER: Script-Normalized WER for Multi-Script Indic ASR EvaluationSometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge DistillationFLEURS-Kobani: Extending the FLEURS Dataset for Northern KurdishIDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector MappingM-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASRScalable Controllable Accented TTSDeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error CorrectionRobust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker CodesCMU's IWSLT 2025 Simultaneous Speech Translation SystemEvaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech RecognitionDysarthria Normalization via Local Lie Group Transformations for Robust
ASRAn Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASRWhistle: Data-Efficient Multilingual and Crosslingual Speech Recognition
via Weakly Phonetic SupervisionAdaptive multilingual speech recognition with pretrained modelsExploring Capabilities of Monolingual Audio Transformers using Large
Datasets in Automatic Speech Recognition of CzechCORAA: a large corpus of spontaneous and prepared speech manually
validated for speech recognition in Brazilian PortugueseSupervised Contrastive Learning for Accented Speech RecognitionAdvancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces
and ConformersTextless Speech-to-Speech Translation With Limited Parallel DataCustom Data Augmentation for low resource ASR using Bark and
Retrieval-Based Voice ConversionGigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and RefinementAsk2Mask: Guided Data Selection for Masked Speech ModelingSpeech Corpora Divergence Based Unsupervised Data Selection for ASRSome voices are too common: Building fair speech recognition systems
using the Common Voice datasetXLSR-Transducer: Streaming ASR for Self-Supervised Pretrained ModelsPseudo-Labeling for Massively Multilingual Speech RecognitionXLS-R: Self-supervised Cross-lingual Speech Representation Learning at
ScaleImproving the transferability of speech separation by meta-learningDistilling a Pretrained Language Model to a Multilingual ASR ModelASR2K: Speech Recognition for Around 2000 Languages without AudioMeWEHV: Mel and Wave Embeddings for Human Voice TasksCan we use Common Voice to train a Multi-Speaker TTS system?Iterative pseudo-forced alignment by acoustic CTC loss for
self-supervised ASR domain adaptationUnsupervised ASR via Cross-Lingual Pseudo-LabelingTranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation
Learning for Cross-lingual Speech RecognitionBoosting End-to-End Multilingual Phoneme Recognition through Exploiting
Universal Speech Attributes ConstraintsConnecting Speech Encoder and Large Language Model for ASRSSHR: Leveraging Self-supervised Hierarchical Representations for
Multilingual Automatic Speech RecognitionLUPET: Incorporating Hierarchical Information Path into Multilingual ASRGLOBE: A High-quality English Corpus with Global Accents for Zero-shot
Speaker Adaptive Text-to-SpeechLow-Resourced Speech Recognition for Iu Mien Language via
Weakly-Supervised Phoneme-based Multilingual Pre-trainingImproving noisy student training for low-resource languages in
End-to-End ASR using CycleGAN and inter-domain lossesLarge Language Model Should Understand Pinyin for Chinese ASR Error
CorrectionFast Streaming Transducer ASR Prototyping via Knowledge Distillation
with WhisperCVSS Corpus and Massively Multilingual Speech-to-Speech TranslationIndonesian Automatic Speech Recognition with XLSR-53A Comparative Analysis of Bilingual and Trilingual Wav2Vec Models for
Automatic Speech Recognition in Multilingual Oral History Archives