Awesome Speech Audio
Papers
Topics
Trending
Leaderboards
Authors
Datasets
Learn
Ask AI
Map
Tools
Packs
News
Videos
All sections
Browse
Papers
The full index, filterable and sortable.
Topics
The same papers grouped by subject tag.
Trending
What moved this week, and by how much.
Authors
Who publishes here, and who they publish with.
Map
The collection laid out by embedding similarity.
Compare
Leaderboards
Benchmark tables, with the paper behind each number.
Datasets
The datasets these papers train and evaluate on.
Tools
Code and libraries released alongside the papers.
Read
Learn
Ordered routes from background reading to current work.
Ask AI
Ask a question and get answers cited to these papers.
Videos
The most-watched talks and lectures in this field.
Packs
Short curated sets built around one question.
Follow
News
Press and coverage tied back to the papers.
Blogs
Author and lab write-ups of their own work.
Newsletter
Email digest of what changed, on a schedule.
Research Radar
Paste an abstract, get matches across every collection.
Yours
Saved
Papers you bookmarked in this browser.
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
Ask AI
Zhijie Yan — most-cited papers & profile · Speech Audio
← authors
·
overview
Zhijie Yan
12
papers ·
103
citations ·
24
h-index
Beijing Sport University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
2025 · 86 citations
A Real-time Speaker Diarization System Based on Spatial Spectrum
2021 · 17 citations
OmniAudio: Generating Spatial Audio from 360-Degree Video
2025
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
2025
Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)
2023
MUG: A General Meeting Understanding and Generation Benchmark
2023
Accurate and Reliable Confidence Estimation Based on Non-Autoregressive End-to-End Speech Recognition System
2023
The second multi-channel multi-party meeting transcription challenge (M2MeT) 2.0): A benchmark for speaker-attributed ASR
2023
Top co-authors
Qian Chen
· 5
Shiliang Zhang
· 5
Wen Wang
· 4
Chong Deng
· 3
Jiaqing Liu
· 3
Qinglin Zhang
· 3
Zhihao Du
· 3
Zhou Zhao
· 3
Fan Yu
· 2
Hai Yu
· 2
Haoneng Luo
· 2
Jinglin Liu
· 2
Topics
cs.SD
eess.AS
cs.CL
cs.AI
cs.HC
cs.CV
cs.LG