Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Anurag Arnab — most-cited papers & profile · Multimodal
← authors
·
overview
Anurag Arnab
25
papers ·
797
citations ·
26
h-index
Google (United States) · Google DeepMind (United Kingdom)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
PaLI-X: On Scaling up a Multilingual Vision and Language Model
2023 · 39 citations
CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
2023 · 4 citations
Seg4diff: Unveiling Open-vocabulary Segmentation In Text-to-image Diffusion Transformers
2025
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
2025
Continual Learning In Vision-language Models Via Aligned Model Merging
2025
VicTR: Video-conditioned Text Representations for Activity Recognition
2023
Top co-authors
Arsha Nagrani
· 2
Cordelia Schmid
· 2
Heeseong Shin
· 2
Paul Hongsuck Seo
· 2
Seungryong Kim
· 2
Sunghwan Hong
· 2
Ahmet İşcen
· 1
AJ Piergiovanni
· 1
Alexander Kolesnikov
· 1
Andreas Peter Steiner
· 1
Anelia Angelova
· 1
Austin Waters
· 1
Topics
Vision-Language Models
Video-Language
Image-Text Retrieval
Benchmarks
Visual QA & Reasoning