Awesome Large Language Models
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Guanrou Yang — most-cited papers & profile · Large Language Models
← authors
·
overview
Guanrou Yang
18
papers ·
91
citations ·
5
h-index
Shanghai Jiao Tong University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
2025 · 86 citations
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
2024 · 4 citations
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
2025 · 1 citations
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
2026
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
2026
SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents
2026
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
2026
DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
2025
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
2025
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
2025
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
2025
Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation
2023
Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning
2023
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
2024
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
2024
Top co-authors
Baosong Yang
· 1
Bin Ma
· 1
Changfeng Gao
· 1
Chong Deng
· 1
Chongjia Ni
· 1
Chong Zhang
· 1
Fan Yu
· 1
Haoneng Luo
· 1
Hao Wang
· 1
Hui Wang
· 1
Jialong Tang
· 1
Jiaqing Liu
· 1
Topics
Speech Recognition
Text-to-Speech
Audio Generation
eess.AS
Multimodal Audio
cs.CL
Speech Translation
cs.SD
Audio Understanding
Voice Cloning