Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Xiao Xu — most-cited papers & profile · Multimodal
← authors
·
overview
Xiao Xu
20
papers ·
44
citations ·
4
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
2024 · 9 citations
BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning
2022 · 3 citations
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
2026
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
2025
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
2025
Magicvl-2b: Empowering Vision-language Models On Mobile Devices With Lightweight Visual Encoders Via Curriculum Learning
2025
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
2025
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
2024
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
2024
Top co-authors
Wanxiang Che
· 5
Libo Qin
· 4
Min-Yen Kan
· 3
Chenfei Wu
· 2
Qiguang Chen
· 2
Yuxi Xie
· 2
Zhi Chen
· 2
Alex Jinpeng Wang
· 1
Bo Yang
· 1
Chenxu Lv
· 1
Dayiheng Liu
· 1
Gengze Zhou
· 1
Topics
Vision-Language Models
Video-Language
Benchmarks
Visual QA & Reasoning
Embodied & Agents
Image-Text Retrieval
Audio-Visual