Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yuxun Tang — most-cited papers & profile · Speech Audio
← authors
·
overview
Yuxun Tang
17
papers ·
13
citations
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
2024 · 5 citations
SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan
2024 · 3 citations
SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
2024 · 3 citations
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
2023 · 2 citations
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
2026
Robust Training of Singing Voice Synthesis Using Prior and Posterior Uncertainty
2025
Adapting Speech Language Model to Singing Voice Synthesis
2025
SingingSDS: A Singing-Capable Spoken Dialogue System for Conversational Roleplay Applications
2025
CartoonSing: Unifying Human and Nonhuman Timbres in Singing Generation
2025
Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing
2024
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
2024
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
2024
TokSing: Singing Voice Synthesis based on Discrete Tokens
2024
SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
2024
Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
2024
Top co-authors
Shinji Watanabe
· 11
Jiatong Shi
· 10
Jionghao Han
· 8
Jiatong Shi
· 6
Jinchuan Tian
· 4
Haibin Wu
· 3
Jia Qi Yip
· 3
Qin Jin
· 3
William Chen
· 3
Yiwen Zhao
· 3
You Zhang
· 3
Yuning Wu
· 3
Topics
Music Generation
Audio Generation
Speech Recognition
Text-to-Speech
Audio Understanding
Speech Enhancement
Voice Cloning
Multimodal Audio