Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Bryan Catanzaro — most-cited papers & profile · Multimodal
← authors
·
overview
Bryan Catanzaro
65
papers ·
1420
citations ·
46
h-index
Nvidia (United States)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Hierarchical Multi-Scale Attention for Semantic Segmentation
2020 · 347 citations
High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs
2017 · 301 citations
eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
2022 · 223 citations
Video-to-Video Synthesis
2018 · 127 citations
DiffWave: A Versatile Diffusion Model for Audio Synthesis
2020 · 121 citations
Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
2020 · 81 citations
PrefixRL: Optimization of Parallel Prefix Circuits using Deep Reinforcement Learning
2022 · 50 citations
BigVGAN: A Universal Neural Vocoder with Large-Scale Training
2022 · 46 citations
Improving Semantic Segmentation via Video Propagation and Label Relaxation
2018 · 37 citations
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
2025 · 24 citations
Image Inpainting for Irregular Holes Using Partial Convolutions
2018 · 18 citations
Can $Q$-Learning with Graph Networks Learn a Generalizable Branching Heuristic for a SAT Solver?
2019 · 11 citations
Compact Language Models via Pruning and Knowledge Distillation
2024 · 7 citations
Introduction to the 1st Place Winning Model of OpenImages Relationship Detection Challenge
2018 · 6 citations
Dual Contrastive Loss and Attention for GANs
2021 · 5 citations
Topics
Training Techniques
Efficiency
Audio Generation
Model Architecture
Text-to-Speech
Fine-Tuning
Reinforcement Learning
Evaluation
Segmentation
Model-Based RL