Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Lei Xie — most-cited papers & profile · Speech Audio
← authors
·
overview
Lei Xie
34
papers ·
137
citations ·
31
h-index
Xi'an Jiaotong University · Nanjing University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
2025 · 51 citations
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
2025 · 19 citations
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
2025 · 10 citations
Efficient Scaling for LLM-based ASR
2025 · 9 citations
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
2025 · 9 citations
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
2025 · 5 citations
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
2025 · 4 citations
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
2025 · 2 citations
HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
2025 · 2 citations
WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
2025 · 1 citations
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
2025 · 1 citations
FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation
2026
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement
2026
LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
2025
DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching
2025
Top co-authors
Chengyou Wang
· 4
Guobin Ma
· 4
Yuepeng Jiang
· 4
Bingshen Mu
· 3
Dake Guo
· 3
Hanke Xie
· 3
Huakang Chen
· 3
Jixun Yao
· 3
Shuai Wang
· 3
Shuiyuan Wang
· 3
Wenjie Tian
· 3
Xuelong Geng
· 3
Topics
Speech Recognition
Audio Generation
Audio Understanding
Speech Enhancement
Text-to-Speech
Speech Translation
Music Generation
Multimodal Audio
Voice Cloning
Speaker Analysis