Awesome Speech Audio
๐
Papers
๐งญ
Topics
๐ฅ
Trending
๐บ๏ธ
Map
๐
Leaderboards
๐
Learn
๐ค
Ask AI
โฏ
More
๐ฅ
Authors
๐
Reading Packs
๐
Datasets
๐ ๏ธ
Tools
๐ฐ
News
๐
Blogs
โ๏ธ
Newsletter
๐ฏ
Research Radar
๐
Saved
+ Add Paper
โพ
โ
โ authors
ยท
overview
Loading authorโฆ
๐ค
Ask AI
Kai Yu โ most-cited papers & profile ยท Speech Audio
โ authors
ยท
overview
Kai Yu
92
papers ยท
256
citations ยท
26
h-index
Shanghai Jiao Tong University ยท Shandong University of Science and Technology
Google Scholar โ
Semantic Scholar โ
OpenAlex โ
Most-cited papers
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
2022 ยท 43 citations
On Modular Training of Neural Acoustics-to-Word Model for LVCSR
2018 ยท 35 citations
DiffVoice: Text-to-Speech with Latent Diffusion
2023 ยท 18 citations
Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
2020 ยท 14 citations
Rich Prosody Diversity Modelling with Phone-level Mixture Density Network
2021 ยท 13 citations
Unsupervised word-level prosody tagging for controllable speech synthesis
2022 ยท 10 citations
Encoder-decoder with Focus-mechanism for Sequence Labelling Based Spoken Language Understanding
2016 ยท 7 citations
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
2025 ยท 5 citations
Diverse and Vivid Sound Generation from Text Descriptions
2023 ยท 4 citations
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
2023 ยท 4 citations
Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation
2023 ยท 4 citations
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
2025 ยท 3 citations
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
2025 ยท 3 citations
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
2024 ยท 3 citations
Deep Discriminant Analysis for i-vector Based Robust Speaker Recognition
2018 ยท 3 citations
Top co-authors
Xie Chen
ยท 33
Shuai Wang
ยท 14
Ziyang Ma
ยท 12
Bohan Li
ยท 10
Haoyu Li
ยท 8
Jing Peng
ยท 7
Yifan Yang
ยท 6
Hao Li
ยท 5
Yanmin Qian
ยท 5
Haoran Wang
ยท 4
Hui Zhang
ยท 4
Junjie Li
ยท 3
Topics
Text-to-Speech
Audio Generation
Speech Recognition
eess.AS
Audio Understanding
cs.SD
cs.AI
Speech Enhancement
Speech Translation
Voice Cloning