Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Alessandro Suglia — most-cited papers & profile · Multimodal
← authors
·
overview
Alessandro Suglia
13
papers ·
5
citations
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
2021 · 5 citations
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
2026
Same Answer, Different Representations: Hidden instability in VLMs
2026
Voyagervision: Investigating The Role Of Multi-modal Information For Open-ended Learning Systems
2025
Multitask Multimodal Prompted Training for Interactive Embodied Task Completion
2023
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
2024
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
2024
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
2024
Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models
2024
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
2024
Top co-authors
Georgios Pantazopoulos
· 5
Malvina Nikandrou
· 5
Arash Eshghi
· 3
Ioannis Konstas
· 3
Amit Parekh
· 2
Oliver Lemon
· 2
Aryo Pradipta Gema
· 1
Bhathiya Hemanthage
· 1
Ethan Smyth
· 1
Fabrizio Silvestri
· 1
Farooq Ahmad Wani
· 1
Fazl Barez
· 1
Topics
Vision-Language Models
Benchmarks
Visual QA & Reasoning
Video-Language
Embodied & Agents
Instruction Tuning
Audio-Visual
Image-Text Retrieval