Awesome Speech Audio
Papers
Topics
Trending
Leaderboards
Authors
Datasets
Learn
Ask AI
Map
Tools
Packs
News
Videos
All sections
Browse
Papers
The full index, filterable and sortable.
Topics
The same papers grouped by subject tag.
Trending
What moved this week, and by how much.
Authors
Who publishes here, and who they publish with.
Map
The collection laid out by embedding similarity.
Compare
Leaderboards
Benchmark tables, with the paper behind each number.
Datasets
The datasets these papers train and evaluate on.
Tools
Code and libraries released alongside the papers.
Read
Learn
Ordered routes from background reading to current work.
Ask AI
Ask a question and get answers cited to these papers.
Videos
The most-watched talks and lectures in this field.
Packs
Short curated sets built around one question.
Follow
News
Press and coverage tied back to the papers.
Blogs
Author and lab write-ups of their own work.
Newsletter
Email digest of what changed, on a schedule.
Research Radar
Paste an abstract, get matches across every collection.
Yours
Saved
Papers you bookmarked in this browser.
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
Ask AI
Yu Zhou — most-cited papers & profile · Speech Audio
← authors
·
overview
Yu Zhou
96
papers ·
74
citations ·
46
h-index
Henan University of Science and Technology · Xiamen University · Nankai University · Xiamen Tobacco Industry (China) · Xiamen University of Technology · Northeastern University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Step-Audio 2 Technical Report
2025 · 1 citations
Semantic Audio-Visual Navigation in Continuous Environments
2026
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
2026
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
2025
Sparse Autoencoder Insights on Voice Embeddings
2025
Source-Critical Reinforcement Learning for Transferring Spoken Language Understanding to a New Language
2018
Memory Consolidation for Contextual Spoken Language Understanding with Dialogue Logistic Inference
2019
Augmenting Slot Values and Contexts for Spoken Language Understanding with Pretrained Models
2021
Look&Listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement
2022
Life-long Learning for Multilingual Neural Machine Translation with Knowledge Distillation
2022
Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection
2023
Self-Modifying State Modeling for Simultaneous Machine Translation
2024
Top co-authors
Jiajun Zhang
· 4
Chengqing Zong
· 3
He Bai
· 2
Lu Xiang
· 2
and Chengqing Zong
· 1
Bingxin Li
· 1
Bin Wang
· 1
Binxing Jiao
· 1
Bo Li
· 1
Boyong Wu
· 1
Brian Li
· 1
Buyun Ma
· 1
Topics
cs.CL
cs.SD
eess.AS
cs.CV
cs.LG
Audio Understanding
Speech Translation
cs.AI
cs.MM