Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 60 papers Β· trending (default) numbers = π₯ heat
Qwen3.5-Omni Technical Report (2026) Qwen Team
7.08 MentalThink: Shaping Thoughts in Mental SVG World (2026) Kangheng Lin et al.
2.00 Ministral 3 (2026) Alexander H. Liu et al.
1.94 UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation (2026) Ruiheng Zhang et al.
1.94 HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding (2026) Haowei Zhang et al.
1.94 Qwen3-ASR Technical Report (2026) Xian Shi et al.
1.94 Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation (2026) Zihan Su et al.
1.94 SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding (2026) Ahmed Y. Radwan et al.
1.94 Enhancing Multi-Image Understanding through Delimiter Token Scaling (2026) Minyoung Lee et al.
1.94 C-ΞΞ: Circuit-Restricted Weight Arithmetic for Selective Refusal (2026) Aditya Kasliwal et al.
1.94 Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders (2026) Boqiang Zhang et al.
1.94 Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models (2026) Linghao Zhang et al.
1.94 VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding (2026) Ruoliu Yang et al.
1.94 VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification (2026) Jiahao Meng et al.
1.94 Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music (2026) Sreyan Ghosh et al.
1.94 MiA-Signature: Approximating Global Activation for Long-Context Understanding (2026) Yuqing Li et al.
1.94 Watch, Remember, Reason: Human-View Video Understanding with MLLMs (2026) Jiahao Meng et al.
1.94 Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning (2026) Jiayi Lei et al.
1.94 UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks (2026) Zhekai Chen et al.
1.94 DocAtlas: Multilingual Document Understanding Across 80+ Languages (2026) Ahmed Heakl et al.
1.89 MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning (2026) Yuxin Liu et al.
1.89 OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding (2026) Ruixiang Zhao et al.
1.89 VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation (2026) Bo Li et al.
1.89 Small Vision-Language Models are Smart Compressors for Long Video Understanding (2026) Junjie Fei et al.
1.83 Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing (2026) Zeyue Tian et al.
1.83 MAIC-UI: Making Interactive Courseware with Generative UI (2026) Shangqing Tu et al.
1.83 HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation (2026) Xin Zhou et al.
1.83 Beyond Language Modeling: An Exploration of Multimodal Pretraining (2026) Shengbang Tong et al.
1.78 Qianfan-OCR: A Unified End-to-End Model for Document Intelligence (2026) Daxiang Dong et al.
1.78 BEAVER: A Training-Free Hierarchical Prompt Compression Method via Structure-Aware Page Selection (2026) Zhengpei Hu et al.
1.78 No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding (2026) Vynska Amalia Permadi et al.
1.72 DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference (2026) Aditya Kumar Singh et al.
1.72 Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of
Images and Videos (2025) Haobo Yuan et al.
1.28 Centurio: On Drivers of Multilingual Ability of Large Vision-Language
Model (2025) Gregor Geigle et al.
1.28 CAD-Editor: A Locate-then-Infill Framework with Automated Training Data
Synthesis for Text-Based CAD Editing (2025) Yu Yuan et al.
1.28 EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language
Models for Vision-Driven Embodied Agents (2025) Rui Yang et al.
1.28 Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue
Learning (2025) Jiazheng Liu et al.
1.28 Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers (2025) Weiming Ren et al.
1.28 VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model (2025) Haozhan Shen et al.
1.28 Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding (2025) Tao Zhang et al.
1.28 Self-alignment of Large Video Language Models with Refined Regularized
Preference Optimization (2025) Pritam Sarkar et al.
1.28 Phi-4-reasoning Technical Report (2025) Marah Abdin et al.
1.28 Unified Multimodal Understanding and Generation Models: Advances,
Challenges, and Opportunities (2025) Xinjie Zhang et al.
1.28 Seed1.5-VL Technical Report (2025) Dong Guo et al.
1.28 Skywork-VL Reward: An Effective Reward Model for Multimodal
Understanding and Reasoning (2025) Xiaokun Wang et al.
1.28 BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture,
Training and Dataset (2025) Jiuhai Chen et al.
1.28 Scaling Computer-Use Grounding via User Interface Decomposition and
Synthesis (2025) Tianbao Xie et al.
1.28 Deep Video Discovery: Agentic Search with Tool Use for Long-form Video
Understanding (2025) Xiaoyi Zhang et al.
1.28 MMPerspective: Do MLLMs Understand Perspective? A Comprehensive
Benchmark for Perspective Perception, Reasoning, and Robustness (2025) Yunlong Tang et al.
1.28 WHEN TO ACT, WHEN TO WAIT: Modeling Structural Trajectories for Intent
Triggerability in Task-Oriented Dialogue (2025) Yaoyao Qian et al.
1.28 OmniGen2: Exploration to Advanced Multimodal Generation (2025) Chenyuan Wu et al.
1.28 Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation
from Diffusion Models (2025) Tiezheng Zhang et al.
1.28 OST-Bench: Evaluating the Capabilities of MLLMs in Online
Spatio-temporal Scene Understanding (2025) JingLi Lin et al.
1.28 Sound and Complete Neuro-symbolic Reasoning with LLM-Grounded
Interpretations (2025) Bradley P. Allen et al.
1.28 Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding
and Generation (2025) Peiyu Wang et al.
1.28 Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with
Long-Term Memory (2025) Lin Long et al.
1.28 AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs (2025) Sidharth Surapaneni et al.
1.28 When Big Models Train Small Ones: Label-Free Model Parity Alignment for
Efficient Visual Question Answering using Small VLMs (2025) Abhirama Subramanyam Penamakuri et al.
1.28 GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric
Reasoning (2025) Guizhen Chen et al.
1.28 UniPixel: Unified Object Referring and Segmentation for Pixel-Level
Visual Reasoning (2025) Ye Liu et al.
1.28