Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Wei Liu — most-cited papers & profile · Multimodal
← authors
·
overview
Wei Liu
298
papers ·
3589
citations ·
111
h-index
Leiden University · Rensselaer Polytechnic Institute · Chinese Academy of Medical Sciences & Peking Union Medical College · Xi'an Polytechnic University · Huazhong University of Science and Technology · Chongqing University of Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Learning Modality Interaction for Temporal Sentence Localization and Event Captioning in Videos
2020 · 11 citations
Seedream 4.0: Toward Next-generation Multimodal Image Generation
2025 · 1 citations
Finebadminton: A Multi-level Dataset For Fine-grained Badminton Video Understanding
2025 · 1 citations
Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
2026
Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
2025
V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
2026
InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward
2026
MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine
2026
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
2026
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
2026
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
2025
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
2025
A Survey on Video Temporal Grounding with Multimodal Large Language Model
2025
MiMo-VL Technical Report
2025
X-driver: Explainable Autonomous Driving With Vision-language Models
2025
Top co-authors
Chong Ma
· 2
Ding Liu
· 2
Hanming Deng
· 2
Haozhe Liu
· 2
Hongfa Wang
· 2
Jie Jiang
· 2
Jie Yang
· 2
J\"urgen Schmidhuber
· 2
Peng Wang
· 2
Sen He
· 2
Shengnan Ma
· 2
Tao Xiang
· 2
Topics
Vision-Language Models
Benchmarks
Video-Language
Visual QA & Reasoning
Embodied & Agents
Image-Text Retrieval
Instruction Tuning