Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yu Gu — most-cited papers & profile · Speech Audio
← authors
·
overview
Yu Gu
54
papers ·
231
citations ·
22
h-index
University of Science and Technology of China · Hong Kong Polytechnic University · Northeastern University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
ByteSing: A Chinese Singing Voice Synthesis System Using Duration Allocated Encoder-Decoder Acoustic Models and WaveRNN Vocoders
2020 · 17 citations
Multi-task WaveNet: A Multi-task Generative Model for Statistical Parametric Speech Synthesis without Fundamental Frequency Conditions
2018 · 5 citations
Rep2wav: Noise Robust text-to-speech Using self-supervised representations
2023 · 1 citations
DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis
2023 · 1 citations
HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis
2026
JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
2025
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
2025
USM-VC: Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion
2025
Waveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension
2018
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
2024
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
2024
LCM-SVC: Latent Diffusion Model Based Singing Voice Conversion with Inference Acceleration via Latent Consistency Distillation
2024
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
2024
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
2024
CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder
2024
Top co-authors
Jie Zhang
· 6
Lirong Dai
· 6
Shihao Chen
· 3
Dan Su
· 2
Na Li
· 2
Yuchen Hu
· 2
Yuxuan Wang
· 2
Zhen-Hua Ling
· 2
Fan Yu
· 1
Jun Wang
· 1
Junxi Liu
· 1
Lin Hu
· 1
Topics
Audio Generation
Text-to-Speech
Music Generation
Voice Cloning
Speech Enhancement
Multimodal Audio
cs.SD
eess.AS
Speaker Analysis
Speech Recognition