Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yi Zhao — most-cited papers & profile · Speech Audio
← authors
·
overview
Yi Zhao
44
papers ·
681
citations ·
0
h-index
Uppsala University · The University of Sydney · Shenzhen University · Xinjiang Institute of Engineering · Faculty of Design · Xi'an Jiaotong University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Voice Conversion Challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion
2020 · 34 citations
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
2025 · 19 citations
MELONS: generating melody with long-term structure using transformers and structure graph
2021 · 7 citations
Improved Prosody from Learned F0 Codebook Representations for VQ-VAE Speech Waveform Reconstruction
2020 · 3 citations
Wasserstein GAN and Waveform Loss-based Acoustic Model Training for Multi-speaker Text-to-Speech Synthesis Systems Using a WaveNet Vocoder
2018 · 1 citations
Transferring neural speech waveform synthesizers to musical instrument sounds generation
2019 · 1 citations
Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis
2020 · 1 citations
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
2024
Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm
2020
High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models
2023
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
2024
Top co-authors
Xin Wang
· 2
Cheng Gong
· 1
Chen Zhang
· 1
Daisuke Saito
· 1
Dake Guo
· 1
Fengrun Zhang
· 1
Hao Li
· 1
Haoyu Li
· 1
Jianhua Tao
· 1
Lei Xie
· 1
Nobuaki Minematsu
· 1
Ran Zhang
· 1
Topics
Audio Generation
Text-to-Speech
Voice Cloning
Music Generation
Speech Recognition
Speaker Analysis
cs.SD
Speech Translation
Multimodal Audio
Speech Enhancement