Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Zhehuai Chen — most-cited papers & profile · Speech Audio
← authors
·
overview
Zhehuai Chen
30
papers ·
296
citations ·
17
h-index
Nvidia (United States)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
2023 · 112 citations
MAESTRO: Matched Speech Text Representations through Modality Matching
2022 · 71 citations
On Modular Training of Neural Acoustics-to-Word Model for LVCSR
2018 · 35 citations
Injecting Text in Self-Supervised Speech Pretraining
2021 · 25 citations
Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
2020 · 14 citations
Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
2022 · 12 citations
Unsupervised Data Selection via Discrete Speech Representation for ASR
2022 · 9 citations
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
2024 · 6 citations
Understanding Shared Speech-Text Representations
2023 · 5 citations
Linguistic Search Optimization for Deep Learning Based LVCSR
2018 · 2 citations
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
2024 · 2 citations
End-to-end contextual speech recognition using class language models and a token passing decoder
2018 · 1 citations
JOIST: A Joint Speech and Text Streaming Model For ASR
2022 · 1 citations
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
2024 · 1 citations
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
2025
Top co-authors
Boris Ginsburg
· 12
Jagadeesh Balam
· 10
Yu Zhang
· 8
Andrew Rosenberg
· 7
Bhuvana Ramabhadran
· 7
Ankur Bapna
· 6
Gary Wang
· 6
Piotr \.Zelasko
· 6
Chao-Han Huck Yang
· 5
Oleksii Hrinchuk
· 5
He Huang
· 4
Ke Hu
· 4
Topics
Speech Recognition
Text-to-Speech
Speech Translation
Multimodal Audio
Audio Understanding
Speech Enhancement
Audio Generation
Music Generation
Speaker Analysis