Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
Vision
loadingβ¦
π€
Ask AI
Awesome Vision β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
Vision
16 papers tagged Vision β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
16 papers Β· trending (default)
numbers = π₯ heat
Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentation
(2025)
Xin Zhang et al.
6.34
MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
(2026)
Honglin Lin et al.
1.94
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
(2026)
Boqiang Zhang et al.
1.94
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm
(2026)
Karan Goyal
1.94
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
(2026)
Guozhen Zhang et al.
1.94
Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning
(2026)
Jiayi Lei et al.
1.94
Introducing Visual Perception Token into Multimodal Large Language Model
(2025)
Runpeng Yu et al.
1.28
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
(2025)
Ahmed Nassar et al.
1.28
EfficientLLM: Efficiency in Large Language Models
(2025)
Zhengqing Yuan et al.
1.28
ImgEdit: A Unified Image Editing Dataset and Benchmark
(2025)
Yang Ye et al.
1.28
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
(2025)
Chengqi Duan et al.
1.28
HunyuanOCR Technical Report
(2025)
Hunyuan Vision Team et al.
1.28
LocalMamba: Visual State Space Model with Windowed Selective Scan
(2024)
Tao Huang et al.
β
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
(2024)
Xiaoyi Dong et al.
β
Mitigating Object Hallucination via Concentric Causal Attention
(2024)
Yun Xing et al.
β
Collaborative Instance Navigation: Leveraging Agent Self-Dialogue to Minimize User Input
(2024)
Francesco Taioli et al.
β