Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yu Zhang — most-cited papers & profile · Speech Audio
← authors
·
overview
Yu Zhang
138
papers ·
4526
citations ·
56
h-index
Southern University of Science and Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
2019 · 3538 citations
A comparison of end-to-end models for long-form speech recognition
2019 · 85 citations
PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS
2021 · 61 citations
Learning Latent Representations for Speech Generation and Transformation
2017 · 53 citations
SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
2021 · 50 citations
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
2019 · 23 citations
Generating diverse and natural text-to-speech samples using a quantized fine-grained VAE and auto-regressive prosody prior
2020 · 15 citations
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
2025 · 12 citations
Speech Recognition with Augmented Synthesized Speech
2019 · 4 citations
Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction
2021 · 2 citations
STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation
2025
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
2025
Echo State Speech Recognition
2021
Pushing the Limits of Non-Autoregressive Speech Recognition
2021
BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
2021
Top co-authors
Chung-Cheng Chiu
· 4
Ye Jia
· 4
Yonghui Wu
· 4
Heiga Zen
· 3
William Chan
· 3
Bhuvana Ramabhadran
· 2
Bo Li
· 2
Daniel S. Park
· 2
Khe Chai Sim
· 2
Quoc V. Le
· 2
Ron J. Weiss
· 2
Ruiqi Li
· 2
Topics
Speech Recognition
Speech Translation
Audio Generation
Text-to-Speech
Music Generation
Multimodal Audio
Audio Understanding
Voice Cloning
Speaker Analysis