Voxpopuli
Emerging14papers using it
23,537HF downloads
157HF likes
2021first seen
Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these
π€ Hugging Faceβ cc0-1.0
Papers using Voxpopuli (14)
- Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabsSpeech Vecalign: an Embedding-based Method for Aligning Parallel Speech DocumentsCMU's IWSLT 2025 Simultaneous Speech Translation SystemOn the use of Performer and Agent Attention for Spoken Language
IdentificationMAESTRO: Matched Speech Text Representations through Modality MatchingExploring Capabilities of Monolingual Audio Transformers using Large
Datasets in Automatic Speech Recognition of CzechMu$^{2}$SLAM: Multitask, Multilingual Speech and Language ModelsSpeechMatrix: A Large-Scale Mined Corpus of Multilingual
Speech-to-Speech TranslationsJoint Pre-Training with Speech and Bilingual Text for Direct Speech to
Speech TranslationPseudo-Labeling for Massively Multilingual Speech RecognitionXLS-R: Self-supervised Cross-lingual Speech Representation Learning at
ScaleTowards Personalization of CTC Speech Recognition Models with Contextual
Adapters and Adaptive BoostingDynamic Chunk Convolution for Unified Streaming and Non-Streaming
Conformer ASRDENOASR: Debiasing ASRs through Selective Denoising