Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yu Zhang — most-cited papers & profile · Speech Audio
← authors
·
overview
Yu Zhang
26
papers ·
4867
citations ·
56
h-index
Southern University of Science and Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Conformer: Convolution-augmented Transformer for Speech Recognition
2020 · 2748 citations
Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
2018 · 475 citations
Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
2018 · 435 citations
W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
2021 · 262 citations
ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context
2020 · 257 citations
A Streaming On-Device End-to-End Model Surpassing Server-Side Conventional Model Quality and Latency
2020 · 202 citations
Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis
2020 · 93 citations
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
2021 · 75 citations
Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling
2020 · 73 citations
Hierarchical Generative Modeling for Controllable Speech Synthesis
2018 · 45 citations
WaveGrad: Estimating Gradients for Waveform Generation
2020 · 44 citations
Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
2019 · 26 citations
Injecting Text in Self-Supervised Speech Pretraining
2021 · 25 citations
Parallel Tacotron: Non-Autoregressive and Controllable TTS
2020 · 21 citations
ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
2019 · 12 citations
Top co-authors
Yonghui Wu
· 12
Ruoming Pang
· 8
Chung-Cheng Chiu
· 7
Heiga Zen
· 6
Ye Jia
· 6
Ron J. Weiss
· 5
Wei Han
· 5
Bo Li
· 4
James Qin
· 4
Jonathan Shen
· 4
Liangliang Cao
· 4
Zhifeng Chen
· 4
Topics
Speech Recognition
Text-to-Speech
Audio Generation
Speech Translation
Music Generation
Audio Understanding
Voice Cloning
Speech Enhancement
Speaker Analysis
Multimodal Audio