Awesome Federated Learning
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
multimodal
loadingβ¦
π€
Ask AI
Awesome multimodal β curated papers, datasets & benchmarks Β· Awesome Federated Learning
β all topics
overview
multimodal
24 papers tagged multimodal β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
24 papers Β· trending (default)
numbers = π₯ heat
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
(2026)
Haoxiang Ma et al.
2.00
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
(2026)
Hohin Kwan et al.
1.94
DataComp-VLM: Improved Open Datasets for Vision-Language Models
(2026)
Matteo Farina et al.
1.94
Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation
(2026)
Jongoh Jeong et al.
1.94
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
(2026)
Yuxuan Fan et al.
1.94
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
(2026)
Chonghuinan Wang et al.
1.94
DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model
(2026)
Daniyel Ayupov et al.
1.94
BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
(2026)
Haitao Wu et al.
1.94
Orca: The World is in Your Mind
(2026)
Yihao Wang et al.
1.94
AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
(2026)
Kien T. Pham et al.
1.94
Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
(2026)
Junha Jung et al.
1.94
ASPIRE: Agentic /Skills Discovery for Robotics
(2026)
Runyu Lu et al.
1.94
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
(2026)
Yoonhyung Park et al.
1.94
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning
(2026)
Hongxing Li et al.
1.94
Attending to Multimodal Generation One Token at a Time
(2026)
Varun Gupta et al.
1.94
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
(2026)
Niu Lian et al.
1.94
Unified Audio Intelligence Without Regressing on Text Intelligence
(2026)
Zhifeng Kong et al.
1.94
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
(2026)
Hairui Zhu et al.
1.94
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
(2026)
Chang Nie et al.
1.94
Vision as Unified Multimodal Generation
(2026)
Xiaoyang Han et al.
1.94
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence
(2026)
Xiangyu Han et al.
1.94
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
(2026)
Hongyu Qu et al.
1.94
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
(2026)
Chen Tang et al.
1.94
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
(2026)
Zhekai Chen et al.
1.94