Multilingual LibriSpeech
Canonical12papers using it
2021first seen
Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset
Papers using Multilingual LibriSpeech (12)
- Pantagruel: Unified Self-Supervised Encoders for French Text and SpeechIDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector MappingUnsupervised Data Selection via Discrete Speech Representation for ASRCML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource
LanguagesPrompting Large Language Models with Speech Recognition AbilitiesREBORN: Reinforcement-Learned Boundary Segmentation with Iterative
Training for Unsupervised ASRJoint Unsupervised and Supervised Training for Multilingual ASRXLS-R: Self-supervised Cross-lingual Speech Representation Learning at
ScaleMulti-blank Transducers for Speech RecognitionEnhancing Unsupervised Speech Recognition with Diffusion GANsAnalyzing and Mitigating Inconsistency in Discrete Audio Tokens for
Neural Codec Language ModelsConfigurable Multilingual ASR with Speech Summary Representations