Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Detai Xin — most-cited papers & profile · Speech Audio
← authors
·
overview
Detai Xin
12
papers ·
27
citations ·
8
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
2024 · 20 citations
Coco-Nut: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-based Control
2023 · 5 citations
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
2024 · 2 citations
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
2026
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
2025
Speaking-Rate-Controllable HiFi-GAN Using Feature Interpolation
2022
Mid-attribute speaker generation using optimal-transport-based interpolation of Gaussian mixture models
2022
Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts
2022
Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech
2023
Building speech corpus with diverse voice characteristics for its prompt-based representation
2024
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
2024
Top co-authors
Hiroshi Saruwatari
· 9
Shinnosuke Takamichi
· 7
Aya Watanabe
· 3
Wataru Nakata
· 3
Yuki Saito
· 3
Dongchao Yang
· 2
Dong Yang
· 2
Jinyu Li
· 2
Takaaki Saeki
· 2
Tomoki Koriyama
· 2
Xu Tan
· 2
Yuancheng Wang
· 2
Topics
Text-to-Speech
Audio Generation
Speaker Analysis
Voice Cloning
Speech Enhancement
Speech Recognition
Music Generation
Multimodal Audio