Awesome Speech Audio
๐
Papers
๐งญ
Topics
๐ฅ
Trending
๐บ๏ธ
Map
๐
Leaderboards
๐
Learn
๐ค
Ask AI
โฏ
More
๐ฅ
Authors
๐
Reading Packs
๐
Datasets
๐ ๏ธ
Tools
๐ฐ
News
๐
Blogs
โ๏ธ
Newsletter
๐ฏ
Research Radar
๐
Saved
+ Add Paper
โพ
โ
โ authors
ยท
overview
Loading authorโฆ
๐ค
Ask AI
Sheng Zhao โ most-cited papers & profile ยท Speech Audio
โ authors
ยท
overview
Sheng Zhao
59
papers ยท
1822
citations ยท
34
h-index
Microsoft Research Asia (China) ยท South China University of Technology
Google Scholar โ
Semantic Scholar โ
OpenAlex โ
Most-cited papers
FastSpeech: Fast, Robust and Controllable Text to Speech
2019 ยท 583 citations
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
2020 ยท 514 citations
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
2023 ยท 163 citations
Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability
2020 ยท 97 citations
AdaSpeech: Adaptive Text to Speech for Custom Voice
2021 ยท 79 citations
Almost Unsupervised Text to Speech and Automatic Speech Recognition
2019 ยท 40 citations
Neural Speech Synthesis with Transformer Network
2018 ยท 39 citations
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
2023 ยท 38 citations
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
2022 ยท 35 citations
MultiSpeech: Multi-Speaker Text to Speech with Transformer
2020 ยท 30 citations
Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
2023 ยท 25 citations
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
2024 ยท 25 citations
Semantic Mask for Transformer based End-to-End Speech Recognition
2019 ยท 24 citations
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
2024 ยท 20 citations
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
2023 ยท 14 citations
Top co-authors
Xu Tan
ยท 31
Yanqing Liu
ยท 27
Jinyu Li
ยท 18
Tao Qin
ยท 17
Shujie Liu
ยท 16
Tie-Yan Liu
ยท 12
Long Zhou
ยท 8
Xiaofei Wang
ยท 8
Bohan Li
ยท 7
Jiang Bian
ยท 6
Yi Ren
ยท 6
Furu Wei
ยท 5
Topics
Text-to-Speech
Audio Generation
Speech Recognition
Speech Translation
Voice Cloning
Speech Enhancement
Music Generation
Multimodal Audio
cs.SD
eess.AS