Awesome AI Agents
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Alessandro Suglia β most-cited papers & profile Β· AI Agents
β authors
Β·
overview
Alessandro Suglia
13
papers Β·
5
citations
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
2021 Β· 5 citations
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
2026
Same Answer, Different Representations: Hidden instability in VLMs
2026
Voyagervision: Investigating The Role Of Multi-modal Information For Open-ended Learning Systems
2025
Multitask Multimodal Prompted Training for Interactive Embodied Task Completion
2023
Is Feedback All You Need? Leveraging Natural Language Feedback in Goal-Conditioned Reinforcement Learning
2023
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
2024
Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
2024
Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation
2024
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
2024
Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models
2024
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
2024
Top co-authors
Amit Parekh
Β· 1
Ioannis Konstas
Β· 1
Kareem Al-Hasan
Β· 1
Malvina Nikandrou
Β· 1
Sabrina McCallum
Β· 1
Topics
Vision-Language Models
Benchmarks
Visual QA & Reasoning
Embodied & Agents
Video-Language
Instruction Tuning
cs.CV
cs.AI
Audio-Visual
RLHF & Alignment