Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
segmentation
loadingβ¦
π€
Ask AI
Awesome segmentation β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
segmentation
19 papers tagged segmentation β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
19 papers Β· trending (default)
numbers = π₯ heat
Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation
(2026)
Zihan Su et al.
1.94
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
(2026)
Shang-Jui Ray Kuo et al.
1.94
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
(2025)
Haobo Yuan et al.
1.28
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
(2025)
Jiayi Lei et al.
1.28
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
(2025)
Muzhi Zhu et al.
1.28
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
(2025)
Tao Zhang et al.
1.28
UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation
(2025)
Linshan Wu et al.
1.28
Optimizing Retrieval-Augmented Generation: Analysis of Hyperparameter Impact on Performance and Efficiency
(2025)
Adel Ammar et al.
1.28
ImgEdit: A Unified Image Editing Dataset and Benchmark
(2025)
Yang Ye et al.
1.28
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
(2025)
Muzhi Zhu et al.
1.28
un^2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
(2025)
Yinqi Li et al.
1.28
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
(2025)
Ye Liu et al.
1.28
Video models are zero-shot learners and reasoners
(2025)
ThaddΓ€us Wiedemer et al.
1.28
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
(2025)
Su Ho Han et al.
1.28
RMT: Retentive Networks Meet Vision Transformers
(2023)
Qihang Fan et al.
β
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
(2023)
Haoxiang Wang et al.
β
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
(2024)
Lianghui Zhu et al.
β
Paint by Inpaint: Learning to Add Image Objects by Removing Them First
(2024)
Navve Wasserman et al.
β
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
(2024)
Yushi Hu et al.
β