WenetSpeech
Emerging15papers using it
2022first seen
Wenetspeech is a multilingual speech dataset used to evaluate the performance of speech-to-text models in aligning speech and text representations across different languages.
Papers using WenetSpeech (15)
- PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMsExploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition SystemsLESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild DataContextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and ActivationZipformer: A faster and better encoder for automatic speech recognitionNextformer: A ConvNeXt Augmented Conformer For End-To-End Speech
Recognition3M: Multi-loss, Multi-path and Multi-level Neural Networks for speech
recognitionWenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech
Generation Model BenchmarkAn Empirical Study of Language Model Integration for Transducer based
Speech RecognitionCB-Conformer: Contextual biasing Conformer for biased word recognitionResearch on an improved Conformer end-to-end Speech Recognition Model
with R-Drop StructureCIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech
RecognitionAutoPrep: An Automatic Preprocessing Framework for In-the-Wild Speech
DataCUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based
Streaming ASRUCorrect: An Unsupervised Framework for Automatic Speech Recognition
Error Correction