Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Dake Guo — most-cited papers & profile · Speech Audio
← authors
·
overview
Dake Guo
19
papers ·
25
citations ·
4
h-index
Northwestern Polytechnical University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
2025 · 19 citations
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
2025 · 5 citations
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
2024 · 1 citations
MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion
2026
FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation
2026
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
2026
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement
2026
VoiceSculptor: Your Voice, Designed By You
2026
Qwen3-TTS Technical Report
2026
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
2025
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
2025
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
2025
HiGNN-TTS: Hierarchical Prosody Modeling with Graph Neural Networks for Expressive Long-form TTS
2023
Automatic channel selection and spatial feature integration for multi-channel speech recognition across various array topologies
2023
Text-aware and Context-aware Expressive Audiobook Speech Synthesis
2024
Top co-authors
Lei Xie
· 15
Jingbin Hu
· 5
Linhan Ma
· 5
Liumeng Xue
· 5
Xinfa Zhu
· 5
Wenhao Li
· 4
Qiang Zhang
· 3
Ziyu Zhang
· 3
Haoyu Zhang
· 2
He Wang
· 2
Jie Liu
· 2
Jin Xu
· 2
Topics
eess.AS
cs.SD
Audio Generation
Text-to-Speech
cs.CL
Speech Recognition
Speech Translation
Voice Cloning
Music Generation