Key papers π₯ Trending (default) π Most cited π Newest first π€ A β Z by title 60 papers Β· trending (default) numbers = π₯ heat
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU (2026) Fan Jiang et al.
17.67 Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents (2026) Hanzhang Zhou et al.
15.56 HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone (2026) Simple AI et al.
14.92 JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents (2026) Yunlong Lin et al.
14.57 ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition (2026) Jooyeol Yun et al.
13.10 StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents (2026) Yan Yang et al.
12.91 KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill (2026) Yunxin Li et al.
12.06 Data Pyramid for Embodied Manipulation (2026) Yifan Ye et al.
11.91 From Pixels to States: Rethinking Interactive World Models as Game Engines (2026) Zhen Li et al.
11.64 Wonder: Video World Model Done Better (2026) Jiacong Xu et al.
9.78 Shieldstral (2026) Antonia Calvi et al.
9.12 Hy-Embodied-VLM-1.0: Efficient Physical-World Agents (2026) Ziyi Wang et al.
7.63 LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 (2026) Mobina Kashaniyan et al.
4.95 Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time Transduction (2026) Jiazhen Huang et al.
4.95 TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering (2026) Zhong Ji et al.
4.95 No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation (2026) Feinan Cheng et al.
4.95 PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics (2026) Haocheng Yin et al.
4.95 Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation (2026) Yu Qi et al.
4.95 Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images (2026) Ronak Bhalgami
4.95 Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs (2026) Yuheng Zong et al.
4.95 Robot-Factored World Models via Robot Rendering (2026) Byungjun Kim et al.
4.95 Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features (2026) Dzmitry Malyshau
4.95 Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots (2026) Hung Nguyen et al.
4.95 Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach (2026) Yinghao Hou et al.
4.95 Progressive Multimodal Alignment for Continual Instruction Tuning (2026) Duzhen Zhang et al.
4.95 What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks (2026) Guanhua Ye et al.
4.33 MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents (2026) Kaixin Ma et al.
4.33 Semantic Anchoring for Robotic Action Representations (2026) Yuan Xu et al.
4.33 Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation (2026) Boyu Mi et al.
4.33 EgoRecovery: Acquiring Failure Recovery Ability Through Human Recovery Demonstration (2026) Zuhao Ge et al.
4.33 Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents (2026) Suman Navaratnarajah et al.
4.33 Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework (2026) Lewen Mi et al.
4.33 CausalGate: Causal Importance Distillation for Transformer Module Pruning (2026) Kiran Nair et al.
4.33 Agentic Autoresearch for CT Reconstruction (2026) Andreas Maier et al.
4.33 Robustifying pathology foundation models via fine-tuning (2026) Alexandre Filiot et al.
4.33 Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests (2026) Patrick Rim et al.
4.33 Layering Virtual Try-On (2026) Chun Feng et al.
4.33 DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video (2026) Muhammad Sulthan Adhipradhana et al.
4.33 Learning Sampling Parameters for Diffusion Models (2026) Arisrei Lim et al.
4.33 Markerless Motion Capture in Routine Clinical Upper Limb Assessments: Validity and Insights Beyond Ordinal Scoring (2026) Tim Unger et al.
4.33 LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory (2026) Haobo Wang et al.
4.33 Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation (2026) Anurag Roy et al.
4.33 What Can I Edit? Open-Ended Strategy Discovery and the Emotion Editability Landscape (2026) Qing Li et al.
4.33 DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning (2026) Mengqi Zhang et al.
4.33 Ambient pressure compensation and robust position control of oil-filled electric joint systems for underwater manipulators (2026) Hongrui Wu et al.
4.33 Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding (2026) Tianyi Gao et al.
4.33 Rethinking Expert Training for Model Merging with Prompt Learning (2026) Christos Georgakilas et al.
4.33 MMOE: Modernizing Diffusion Transformers with Efficient Expert Design (2026) Yanhao Jia et al.
4.33 Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns (2026) Hong Chen et al.
4.33 Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation (2026) Zheng Tong et al.
4.33 Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models (2026) Cl\'ement Grisi et al.
4.33 Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA (2026) Carlos Celemin et al.
4.33 Lottery Tickets Are Not Deployment Tickets (2026) Bum Jun Kim
4.33 Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA (2026) Zhongkuan Mao et al.
4.33 ViP-Rig: Visual-Prompted Controllable Rigging (2026) Zihan Qin et al.
4.33 OPLD: On-Policy Latent Distillation for Multimodal Reasoning (2026) Shoutai Zhu et al.
4.33 UniCross: Unified Cross-Skill Dexterous Manipulation Synthesis (2026) Hui Zhang et al.
4.33 When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More (2026) Zhongyuan Wang et al.
4.27 Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture (2026) Chengshuai Yang
3.45 VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness (2026) Niklas Bunzel
3.45