Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Junyi Ao — most-cited papers & profile · Speech Audio
← authors
·
overview
Junyi Ao
15
papers ·
52
citations ·
8
h-index
Chinese University of Hong Kong, Shenzhen
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
2021 · 30 citations
Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
2022 · 13 citations
Multi-View Self-Attention Based Transformer for Speaker Recognition
2021 · 3 citations
The YiTrans End-to-End Speech Translation System for IWSLT 2022 Offline Shared Task
2022 · 3 citations
SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
2022 · 3 citations
Scaling Speech Tokenizers with Diffusion Autoencoders
2026
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
2025
LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT
2022
CoBERT: Self-Supervised Speech Representation Learning Through Code Representation Learning
2022
token2vec: A Joint Self-Supervised Pre-training Framework Using Unpaired Speech and Text
2022
Self-Supervised Acoustic Word Embedding Learning via Correspondence Transformer Encoder
2023
USED: Universal Speaker Extraction and Diarization
2023
The NUS-HLT System for ICASSP2024 ICMC-ASR Grand Challenge
2023
Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks
2024
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
2024
Top co-authors
Haizhou Li
· 10
Long Zhou
· 6
Shujie Liu
· 5
Tom Ko
· 5
Furu Wei
· 4
Jinyu Li
· 4
Meng Ge
· 3
Xianghu Yue
· 3
Zhihua Wei
· 3
Ziqiang Zhang
· 3
Jingru Lin
· 2
Lirong Dai
· 2
Topics
Speech Recognition
Audio Understanding
Text-to-Speech
Multimodal Audio
Speech Translation
Speaker Analysis
Speech Enhancement
Audio Generation