Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Kai Yu — most-cited papers & profile · Speech Audio
← authors
·
overview
Kai Yu
41
papers ·
150
citations ·
26
h-index
Shanghai Jiao Tong University · Shandong University of Science and Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
On Modular Training of Neural Acoustics-to-Word Model for LVCSR
2018 · 35 citations
Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
2020 · 14 citations
Encoder-decoder with Focus-mechanism for Sequence Labelling Based Spoken Language Understanding
2016 · 7 citations
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
2023 · 4 citations
Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation
2023 · 4 citations
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
2025 · 3 citations
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
2024 · 3 citations
Deep Discriminant Analysis for i-vector Based Robust Speaker Recognition
2018 · 3 citations
End-to-End Monaural Multi-speaker ASR System without Pretraining
2018 · 3 citations
End-to-End Speaker-Dependent Voice Activity Detection
2020 · 2 citations
Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
2023 · 2 citations
Jointly Encoding Word Confusion Network and Dialogue Context with BERT for Spoken Language Understanding
2020 · 1 citations
Future Vector Enhanced LSTM Language Model for LVCSR
2020 · 1 citations
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
2024 · 1 citations
Joint decoding method for controllable contextual speech recognition based on Speech LLM
2025
Top co-authors
Chenpeng Du
· 10
Yiwei Guo
· 9
Xie Chen
· 7
Ziyang Ma
· 6
Xie Chen
· 5
Feiyu Shen
· 4
Hankun Wang
· 4
Yanmin Qian
· 4
Baochen Yang
· 3
Su Zhu
· 3
Zhikang Niu
· 3
Hao Li
· 2
Topics
Speech Recognition
Text-to-Speech
Audio Generation
Audio Understanding
Speech Translation
Speech Enhancement
Speaker Analysis
Voice Cloning
Music Generation