Awesome Speech Audio
Papers
Topics
Trending
Leaderboards
Authors
Datasets
Learn
Ask AI
Map
Tools
Packs
News
Videos
All sections
Browse
Papers
The full index, filterable and sortable.
Topics
The same papers grouped by subject tag.
Trending
What moved this week, and by how much.
Authors
Who publishes here, and who they publish with.
Map
The collection laid out by embedding similarity.
Compare
Leaderboards
Benchmark tables, with the paper behind each number.
Datasets
The datasets these papers train and evaluate on.
Tools
Code and libraries released alongside the papers.
Read
Learn
Ordered routes from background reading to current work.
Ask AI
Ask a question and get answers cited to these papers.
Videos
The most-watched talks and lectures in this field.
Packs
Short curated sets built around one question.
Follow
News
Press and coverage tied back to the papers.
Blogs
Author and lab write-ups of their own work.
Newsletter
Email digest of what changed, on a schedule.
Research Radar
Paste an abstract, get matches across every collection.
Yours
Saved
Papers you bookmarked in this browser.
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
Ask AI
Ke Xu — most-cited papers & profile · Speech Audio
← authors
·
overview
Ke Xu
107
papers ·
553
citations ·
51
h-index
Nanjing Normal University · Southeast University · Hubei University of Arts and Science · Hubei University · South China University of Technology · Tsinghua University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
2026
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
2025
ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
2025
Improving Grammatical Error Correction with Machine Translation Pairs
2019
Improving BERT with Syntax-aware Local Attention
2020
Improving Sequence-to-Sequence Pre-training via Sequence Span Rewriting
2021
Voice2Mesh: Cross-Modal 3D Face Model Generation from Voices
2021
AdvSV: An Over-the-Air Adversarial Attack Dataset for Speaker Verification
2023
MIO: A Foundation Model on Multimodal Tokens
2024
PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment
2024
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
2024
Top co-authors
Wangchunshu Zhou
· 4
Furu Wei
· 2
Jiaheng Liu
· 2
Jie Fu
· 2
Tao Ge
· 2
Wenhao Huang
· 2
Canwen Xu
· 1
Changhao Pan
· 1
Chang Mu
· 1
Chao Li
· 1
Chengfang Fang
· 1
Chin-Cheng Hsu
· 1
Topics
cs.CL
cs.SD
cs.AI
cs.LG
eess.AS
cs.CR
Speech Translation
cs.GR
cs.CV