Awesome Speech Audio
๐
Papers
๐งญ
Topics
๐ฅ
Trending
๐บ๏ธ
Map
๐
Leaderboards
๐
Learn
๐ค
Ask AI
โฏ
More
๐ฅ
Authors
๐
Reading Packs
๐
Datasets
๐ ๏ธ
Tools
๐ฐ
News
๐
Blogs
โ๏ธ
Newsletter
๐ฏ
Research Radar
๐
Saved
+ Add Paper
โพ
โ
โ authors
ยท
overview
Loading authorโฆ
๐ค
Ask AI
Sheng Zhao โ most-cited papers & profile ยท Speech Audio
โ authors
ยท
overview
Sheng Zhao
45
papers ยท
1778
citations ยท
34
h-index
Microsoft Research Asia (China) ยท South China University of Technology
Google Scholar โ
Semantic Scholar โ
OpenAlex โ
Most-cited papers
FastSpeech: Fast, Robust and Controllable Text to Speech
2019 ยท 583 citations
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
2020 ยท 514 citations
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
2023 ยท 163 citations
Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability
2020 ยท 97 citations
AdaSpeech: Adaptive Text to Speech for Custom Voice
2021 ยท 79 citations
Almost Unsupervised Text to Speech and Automatic Speech Recognition
2019 ยท 40 citations
Neural Speech Synthesis with Transformer Network
2018 ยท 39 citations
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
2023 ยท 38 citations
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
2022 ยท 35 citations
MultiSpeech: Multi-Speaker Text to Speech with Transformer
2020 ยท 30 citations
Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
2023 ยท 25 citations
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
2024 ยท 25 citations
Semantic Mask for Transformer based End-to-End Speech Recognition
2019 ยท 24 citations
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
2023 ยท 14 citations
Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion
2019 ยท 13 citations
Top co-authors
Tao Qin
ยท 15
Yanqing Liu
ยท 15
Tie-Yan Liu
ยท 12
Xu Tan
ยท 12
Jinyu Li
ยท 11
Shujie Liu
ยท 9
Lei He
ยท 6
Yi Ren
ยท 6
Manthan Thakker
ยท 5
Yichong Leng
ยท 5
Naoyuki Kanda
ยท 4
Sefik Emre Eskimez
ยท 4
Topics
Text-to-Speech
Audio Generation
Speech Recognition
Speech Translation
Voice Cloning
Speech Enhancement
Music Generation
Multimodal Audio
Speaker Analysis
Audio Understanding