Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Dongchao Yang — most-cited papers & profile · Speech Audio
← authors
·
overview
Dongchao Yang
32
papers ·
159
citations ·
14
h-index
Chinese University of Hong Kong
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
MoonCast: High-Quality Zero-Shot Podcast Generation
2025 · 24 citations
Target Confusion in End-to-end Speaker Extraction: Analysis and Approaches
2022 · 22 citations
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
2024 · 20 citations
HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
2023 · 19 citations
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
2025 · 15 citations
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
2023 · 14 citations
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
2025 · 12 citations
Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information
2021 · 11 citations
NADiffuSE: Noise-aware Diffusion-based Model for Speech Enhancement
2023 · 6 citations
InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt
2023 · 5 citations
Speaker-Aware Mixture of Mixtures Training for Weakly Supervised Speaker Extraction
2022 · 3 citations
Make-A-Voice: Unified Voice Synthesis With Discrete Representation
2023 · 2 citations
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
2024 · 2 citations
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
2024 · 2 citations
NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
2022 · 1 citations
Top co-authors
Helen Meng
· 13
Xixin Wu
· 12
Songxiang Liu
· 9
Xu Tan
· 9
Rongjie Huang
· 6
Yuexian Zou
· 6
Zeqian Ju
· 6
Kai Shen
· 5
Xueyuan Chen
· 5
Yuanyuan Wang
· 5
Dingdong Wang
· 4
Sheng Zhao
· 4
Topics
Audio Generation
Text-to-Speech
Speech Recognition
Speech Enhancement
Audio Understanding
Multimodal Audio
Voice Cloning
Music Generation
Speech Translation
cs.SD