Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
VQA
loadingβ¦
π€
Ask AI
Awesome VQA β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
VQA
19 papers tagged VQA β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
19 papers Β· trending (default)
numbers = π₯ heat
AIBench: Evaluating Visual-Logical Consistency in Academic Illustration Generation
(2026)
Zhaohe Liao et al.
1.94
Task-Focused Memorization for Multimodal Agents
(2026)
Tao Zou et al.
1.94
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
(2026)
Nikita Kachaev et al.
1.94
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
(2026)
Dian Zheng et al.
1.89
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
(2026)
Yu Zeng et al.
1.72
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
(2025)
Xiangyu Zhao et al.
1.28
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
(2025)
Muzhi Zhu et al.
1.28
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
(2025)
Md Mohaiminul Islam et al.
1.28
MMSearch-R1: Incentivizing LMMs to Search
(2025)
Jinming Wu et al.
1.28
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
(2025)
Vatsal Agarwal et al.
1.28
MovieCORE: COgnitive REasoning in Movies
(2025)
Gueter Josmy Faure et al.
1.28
When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
(2025)
Abhirama Subramanyam Penamakuri et al.
1.28
MedVLSynther: Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
(2025)
Xiaoke Huang et al.
1.28
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
(2025)
Mengjie Deng et al.
1.28
HunyuanOCR Technical Report
(2025)
Hunyuan Vision Team et al.
1.28
World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
(2025)
Eunsu Kim et al.
1.28
AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
(2025)
Zichuan Lin et al.
1.28
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
(2025)
Atsuyuki Miyai et al.
1.28
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
(2024)
Neelabh Sinha et al.
β