Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Fan Yu — most-cited papers & profile · Speech Audio
← authors
·
overview
Fan Yu
22
papers ·
193
citations ·
10
h-index
Central South University · Beijing University of Posts and Telecommunications · Hainan University · Huawei Technologies (China) · Beijing Normal University · Nanjing University of Aeronautics and Astronautics
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
2025 · 86 citations
Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
2020 · 46 citations
WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
2021 · 16 citations
The Accented English Speech Recognition Challenge 2020: Open Datasets, Tracks, Baselines, Results and Methods
2021 · 16 citations
M2MeT: The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge
2021 · 11 citations
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
2024 · 4 citations
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
2024 · 3 citations
Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
2021 · 2 citations
JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
2025
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
2025
Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge
2022
A Comparative Study on Speaker-attributed Automatic Speech Recognition in Multi-party Meetings
2022
A Comparative Study on Multichannel Speaker-Attributed Automatic Speech Recognition in Multi-party Meetings
2022
CASA-ASR: Context-Aware Speaker-Attributed ASR
2023
BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR
2023
Top co-authors
Shiliang Zhang
· 14
Zhihao Du
· 11
Lei Xie
· 10
Qian Chen
· 7
Zhifu Gao
· 6
Guanrou Yang
· 5
Pengcheng Guo
· 5
Xian Shi
· 5
Yuxuan Wang
· 4
Zhijie Yan
· 4
Bin Ma
· 3
Changfeng Gao
· 3
Topics
Speech Recognition
Text-to-Speech
Speech Translation
Speaker Analysis
Multimodal Audio
Audio Generation
Audio Understanding
Voice Cloning
cs.SD
eess.AS