Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 60 papers Β· trending (default) numbers = π₯ heat
Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language
Models for Domain-Generalized Semantic Segmentation (2025) Xin Zhang et al.
6.34 Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal
Representations (2025) Jeonghyeon Kim et al.
4.47 An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift (2026) Constantinos Karouzos et al.
1.94 Stochastic CHAOS: Why Deterministic Inference Kills, and Distributional Variability Is the Heartbeat of Artifical Cognition (2026) Tanmay Joshi et al.
1.94 Qwen3-ASR Technical Report (2026) Xian Shi et al.
1.94 THINKSAFE: Self-Generated Safety Alignment for Reasoning Models (2026) Seanie Lee et al.
1.94 Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning (2026) Xu Ma et al.
1.94 BeamPERL: Parameter-Efficient RL with Verifiable Rewards Specializes Compact LLMs for Structured Beam Mechanics Reasoning (2026) Tarjei Paule Hage et al.
1.94 NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval (2026) Zhuchenyang Liu et al.
1.94 Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training (2026) Peng Sun et al.
1.94 Qualixar OS: A Universal Operating System for AI Agent Orchestration (2026) Varun Pratap Bhardwaj
1.94 Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values (2026) Haonan Dong et al.
1.94 Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why (2026) Mohammadreza Armandpour et al.
1.94 SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces (2026) Qi Hu et al.
1.94 Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems (2026) Shaoyang Xu et al.
1.94 OPRD: On-Policy Representation Distillation (2026) Shenzhi Yang et al.
1.94 Watch, Remember, Reason: Human-View Video Understanding with MLLMs (2026) Jiahao Meng et al.
1.94 P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning (2026) Yikang Yang et al.
1.94 The Role of Feedback Alignment in Self-Distillation (2026) Semih Kara et al.
1.94 Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models (2026) Malikeh Ehghaghi et al.
1.94 See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents (2026) Siyi Chen et al.
1.94 Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale (2026) Ang Li et al.
1.94 HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining (2026) Juncheng Ma et al.
1.94 StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs (2026) Shaghayegh Kolli et al.
1.94 PrivacyAlign: Contextual Privacy Alignment for LLM Agents (2026) Manveer Singh Tamber et al.
1.94 A Gravitational Interpretation of Fine-Tuning Reversion (2026) Samuele Poppi et al.
1.94 GEAR: Guided End-to-End AutoRegression for Image Synthesis (2026) Bin Lin et al.
1.94 Dynin-Omni: Omnimodal Unified Large Diffusion Language Model (2026) Jaeik Kim et al.
1.83 MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale (2026) Bin Wang et al.
1.83 Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models (2026) Yunkai Zhang et al.
1.83 LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories (2026) Zhanhao Liang et al.
1.83 LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation (2026) Jiazheng Xing et al.
1.78 Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers (2026) Bozhou Li et al.
1.72 Reliable and Responsible Foundation Models: A Comprehensive Survey (2026) Xinyu Yang et al.
1.72 Data Science and Technology Towards AGI Part I: Tiered Data Management (2026) Yudong Wang et al.
1.72 GeoMotionGPT: Geometry-Aligned Motion Understanding with Large Language Models (2026) Zhankai Ye et al.
1.67 OVD: On-policy Verbal Distillation (2026) Jing Xiong et al.
1.67 Baichuan-Omni-1.5 Technical Report (2025) Yadong Li et al.
1.28 YINYANG-ALIGN: Benchmarking Contradictory Objectives and Proposing
Multi-Objective Optimization based DPO for Text-to-Image Alignment (2025) Amitava Das et al.
1.28 Ask in Any Modality: A Comprehensive Survey on Multimodal
Retrieval-Augmented Generation (2025) Mohammad Mahdi Abootorabi et al.
1.28 Rethinking Diverse Human Preference Learning through Principal Component
Analysis (2025) Feng Luo et al.
1.28 OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference (2025) Xiangyu Zhao et al.
1.28 Multimodal Representation Alignment for Image Generation: Text-Image
Interleaved Control Is Easier Than You Think (2025) Liang Chen et al.
1.28 Q-Eval-100K: Evaluating Visual Quality and Alignment Level for
Text-to-Vision Content (2025) Zicheng Zhang et al.
1.28 ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question
Generation and Answering (2025) Kaisi Guan et al.
1.28 When Words Outperform Vision: VLMs Can Self-Improve Via Text-Only
Training For Human-Centered Decision Making (2025) Zhe Hu et al.
1.28 Beyond Words: Advancing Long-Text Image Generation via Multimodal
Autoregressive Models (2025) Alex Jinpeng Wang et al.
1.28 Skywork-VL Reward: An Effective Reward Model for Multimodal
Understanding and Reasoning (2025) Xiaokun Wang et al.
1.28 Seeing is Believing, but How Much? A Comprehensive Analysis of
Verbalized Calibration in Vision-Language Models (2025) Weihao Xuan et al.
1.28 Adversarial Attacks against Closed-Source MLLMs via Feature Optimal
Alignment (2025) Xiaojun Jia et al.
1.28 OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation (2025) Jingjing Chang et al.
1.28 CAMS: A CityGPT-Powered Agentic Framework for Urban Human Mobility
Simulation (2025) Yuwei Du et al.
1.28 Kwai Keye-VL Technical Report (2025) Kwai Keye Team et al.
1.28 T-LoRA: Single Image Diffusion Model Customization Without Overfitting (2025) Vera Soboleva et al.
1.28 Towards Multimodal Understanding via Stable Diffusion as a Task-Aware
Feature Extractor (2025) Vatsal Agarwal et al.
1.28 Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved
Image Generation (2025) Junyan Ye et al.
1.28 VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use (2025) Dongfu Jiang et al.
1.28 FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning
Dataset and Comprehensive Benchmark (2025) Rongyao Fang et al.
1.28 UniPixel: Unified Object Referring and Segmentation for Pixel-Level
Visual Reasoning (2025) Ye Liu et al.
1.28 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening,
Speaking, and Viewing (2025) Ke Wang et al.
1.28