Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Brian Yan — most-cited papers & profile · Speech Audio
← authors
·
overview
Brian Yan
14
papers ·
55
citations ·
15
h-index
Western University · London Health Sciences Centre
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding
2022 · 28 citations
Improving Massively Multilingual ASR With Auxiliary CTC Objectives
2023 · 26 citations
CTC Alignments Improve Autoregressive Translation
2022 · 1 citations
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
2025
Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
2021
Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation
2022
Align, Write, Re-order: Explainable End-to-End Speech Translation via Operation Sequence Generation
2022
Avoid Overthinking in Self-Supervised Models for Speech Recognition
2022
4D ASR: Joint modeling of CTC, Attention, Transducer, and Mask-Predict decoders
2022
ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit
2023
Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task Generalization
2023
Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning
2023
Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
2023
Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing
2023
Top co-authors
Shinji Watanabe
· 14
Dan Berrebbi
· 5
Jiatong Shi
· 4
Siddharth Dalmia
· 4
Xuankai Chang
· 4
Soumi Maiti
· 3
Yifan Peng
· 3
Yuya Fujita
· 3
Puyuan Peng
· 2
Samuele Cornell
· 2
Wangyou Zhang
· 2
William Chen
· 2
Topics
Speech Recognition
Speech Translation
Audio Understanding
Text-to-Speech
Multimodal Audio
Speech Enhancement
Audio Generation
Music Generation