Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Albert Gatt — most-cited papers & profile · Multimodal
← authors
·
overview
Albert Gatt
8
papers ·
10
citations ·
29
h-index
Utrecht University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
What Vision-Language Models `See' when they See Scenes
2021 · 8 citations
Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
2025
From Image Captioning to Visual Storytelling
2025
Understanding Cross-modal Interactions in V&L Models that Generate Scene Descriptions
2022
CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding
2024
Top co-authors
Kees van Deemter
· 2
Michele Cafagna
· 2
Admitos Passadakis
· 1
Ayman Santeer
· 1
Daniil Ignatev
· 1
Denis Paperno
· 1
Ivana Be\v{n}ov\'a
· 1
Michal Gregor
· 1
Yingjin Song
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Video-Language
Image-Text Retrieval
Benchmarks