Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jiatong Shi — most-cited papers & profile · Speech Audio
← authors
·
overview
Jiatong Shi
32
papers ·
46
citations ·
0
h-index
Carnegie Mellon University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Improving Massively Multilingual ASR With Auxiliary CTC Objectives
2023 · 26 citations
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
2024 · 5 citations
The Singing Voice Conversion Challenge 2023
2023 · 4 citations
SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities
2022 · 3 citations
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
2023 · 2 citations
Improving Cascaded Unsupervised Speech Translation with Denoising Back-translation
2023 · 2 citations
Bridging Speech and Textual Pre-trained Models with Unsupervised ASR
2022 · 1 citations
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
2023 · 1 citations
Exploration on HuBERT with Multiple Resolutions
2023 · 1 citations
Preference Alignment Improves Language Model-Based TTS
2024 · 1 citations
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
2026
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
2026
Robust Training of Singing Voice Synthesis Using Prior and Posterior Uncertainty
2025
CartoonSing: Unifying Human and Nonhuman Timbres in Singing Generation
2025
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
2025
Top co-authors
Shinji Watanabe
· 26
Hung-yi Lee
· 9
Yuxun Tang
· 6
Jinchuan Tian
· 5
William Chen
· 5
Xuankai Chang
· 5
Abdelrahman Mohamed
· 4
Brian Yan
· 4
Dan Berrebbi
· 4
Hirofumi Inaguma
· 4
Shang-Wen Li
· 4
Siddhant Arora
· 4
Topics
Speech Recognition
Audio Generation
Audio Understanding
Music Generation
Speech Translation
Text-to-Speech
Multimodal Audio
Speech Enhancement
Voice Cloning
Speaker Analysis