Awesome Speech Audio
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Bryan Catanzaro — most-cited papers & profile · Speech Audio
← authors
·
overview
Bryan Catanzaro
65
papers ·
1420
citations ·
46
h-index
Nvidia (United States)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
DiffWave: A Versatile Diffusion Model for Audio Synthesis
2020 · 121 citations
Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
2020 · 81 citations
BigVGAN: A Universal Neural Vocoder with Large-Scale Training
2022 · 46 citations
VANI: Very-lightweight Accent-controllable TTS for Native and Non-native speakers with Identity Preservation
2023 · 2 citations
Generative Modeling for Low Dimensional Speech Attributes with Neural Spline Flows
2022 · 1 citations
Multilingual Multiaccented Multispeaker TTS with RADTTS
2023 · 1 citations
PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
2026
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
2025
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
2024
A2SB: Audio-to-Audio Schrodinger Bridges
2025
WaveGlow: A Flow-based Generative Network for Speech Synthesis
2018
Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens
2019
One TTS Alignment To Rule Them All
2021
Speech Denoising in the Waveform Domain with Self-Attention
2022
CleanUNet 2: A Hybrid Speech Denoising Model on Waveform and Spectrogram
2023
Top co-authors
Rafael Valle
· 11
Rohan Badlani
· 5
Wei Ping
· 5
Zhifeng Kong
· 5
Sang-gil Lee
· 4
Ryan Prenger
· 3
Ambrish Dantrey
· 2
Boris Ginsburg
· 2
Sungwon Kim
· 2
Ambuj Mehrish
· 1
Amir Ali Bagherzadeh
· 1
Ante Jukić
· 1
Topics
Audio Generation
Text-to-Speech
Music Generation
Voice Cloning
Speech Recognition
Speech Enhancement
cs.CL
Multimodal Audio
Audio Understanding