Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Xu Tan — most-cited papers & profile · Speech Audio
← authors
·
overview
Xu Tan
19
papers ·
113
citations ·
15
h-index
Beijing University of Civil Engineering and Architecture
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
2023 · 38 citations
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
2022 · 35 citations
PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior
2021 · 23 citations
ResGrad: Residual Denoising Diffusion Probabilistic Models for Text to Speech
2022 · 5 citations
Transformer-S2A: Robust and Efficient Speech-to-Animation
2021 · 2 citations
PromptTTS: Controllable Text-to-Speech with Text Descriptions
2022 · 2 citations
WuYun: Exploring hierarchical skeleton-guided melody generation using knowledge-enhanced deep learning
2023 · 2 citations
FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model
2023 · 2 citations
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
2024 · 2 citations
Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
2022 · 1 citations
DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
2022 · 1 citations
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
2025
Cross-domain Speech Recognition with Unsupervised Character-level Distribution Matching
2021
A study on the efficacy of model pre-training in developing neural text-to-speech system
2021
AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios
2022
Top co-authors
Sheng Zhao
· 12
Lei He
· 6
Tao Qin
· 6
Yanqing Liu
· 5
Yichong Leng
· 5
Yihan Wu
· 4
Haohe Liu
· 3
Tie-Yan Liu
· 3
Daxin Tan
· 2
Detai Xin
· 2
Guangyan Zhang
· 2
Hiroshi Saruwatari
· 2
Topics
Audio Generation
Text-to-Speech
Speech Recognition
Speech Translation
Speech Enhancement
Music Generation
Multimodal Audio
Voice Cloning
Speaker Analysis