Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
object
loadingβ¦
π€
Ask AI
Awesome object β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
object
24 papers tagged object β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
24 papers Β· trending (default)
numbers = π₯ heat
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
(2026)
Lingfeng Ren et al.
1.94
ActWorld: From Explorable to Interactive World Model via Action-Aware Memory
(2026)
Zhexiao Xiong et al.
1.94
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
(2025)
Haobo Yuan et al.
1.28
V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multi-Modal Large Language Models
(2025)
Hsu-kuang Chiu et al.
1.28
Visual-RFT: Visual Reinforcement Fine-Tuning
(2025)
Ziyu Liu et al.
1.28
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
(2025)
Xinhao Li et al.
1.28
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
(2025)
Tao Zhang et al.
1.28
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
(2025)
Muzhi Zhu et al.
1.28
OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
(2025)
Yuan Gong et al.
1.28
Visual Representation Alignment for Multimodal Large Language Models
(2025)
Heeji Yoon et al.
1.28
Video models are zero-shot learners and reasoners
(2025)
ThaddΓ€us Wiedemer et al.
1.28
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
(2025)
Kai-Po Chang et al.
1.28
Tiny LVLM-eHub: Early Multimodal Experiments with Bard
(2023)
Wenqi Shao et al.
β
RMT: Retentive Networks Meet Vision Transformers
(2023)
Qihang Fan et al.
β
A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation
(2023)
Eyal Segalis et al.
β
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
(2024)
Lianghui Zhu et al.
β
Paint by Inpaint: Learning to Add Image Objects by Removing Them First
(2024)
Navve Wasserman et al.
β
FRAP: Faithful and Realistic Text-to-Image Generation with Adaptive Prompt Weighting
(2024)
Liyao Jiang et al.
β
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
(2024)
Yang Liu et al.
β
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
(2024)
Hang Hua et al.
β
Mitigating Object Hallucination via Concentric Causal Attention
(2024)
Yun Xing et al.
β
SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE
(2024)
Yongwei Chen et al.
β
EMOv2: Pushing 5M Vision Model Frontier
(2024)
Jiangning Zhang et al.
β
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
(2024)
Yuqian Yuan et al.
β