CoVoST 2
Emerging28papers using it
2021first seen
CoVoST-2 is a multilingual speech-to-text translation dataset used to evaluate machine translation systems across multiple languages.
Papers using CoVoST 2 (28)
- Prepending or Cross-Attention for Speech-to-Text? An Empirical
ComparisonSpeech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and TranslationScalable Multilingual Multimodal Machine Translation with Speech-Text FusionPART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMsSASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech TranslationSpeech Translation Refinement using Large Language ModelsMAESTRO: Matched Speech Text Representations through Modality MatchingSLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text
Joint Pre-TrainingPre-training for Speech Translation: CTC Meets Optimal TransportSimple and Effective Unsupervised Speech TranslationImproving Cascaded Unsupervised Speech Translation with Denoising
Back-translationComSL: A Composite Speech-Language Model for End-to-End Speech-to-Text
TranslationCompact Speech Translation Models via Discrete Speech Units PretrainingMake More of Your Data: Minimal Effort Data Augmentation for Automatic
Speech Recognition and TranslationTowards a Deep Understanding of Multilingual End-to-End Speech
TranslationLLaST: Improved End-to-end Speech Translation System Leveraged by Large
Language ModelsCTC-GMM: CTC guided modality matching for fast and accurate streaming
speech translationXLS-R: Self-supervised Cross-lingual Speech Representation Learning at
ScaleImproved Cross-Lingual Transfer Learning For Automatic Speech
TranslationImproving End-to-End Speech Translation by Imitation-Based Knowledge
Distillation with Synthetic TranscriptsGenTranslate: Large Language Models are Generative Multilingual Speech
and Machine TranslatorsInvestigating Decoder-only Large Language Models for Speech-to-text
TranslationTask Arithmetic for Language Expansion in Speech TranslationMaking LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum LearningRepresentation Purification for End-to-End Speech TranslationZero-resource Speech Translation and Recognition with LLMsCVSS Corpus and Massively Multilingual Speech-to-Speech TranslationSample, Translate, Recombine: Leveraging Audio Alignments for Data
Augmentation in End-to-end Speech Translation