Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Ran Xu — most-cited papers & profile · Multimodal
← authors
·
overview
Ran Xu
61
papers ·
332
citations ·
0
h-index
Jilin University · Middle East College
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
TAG: Boosting Text-VQA via Text-aware Visual Question-answer Generation
2022 · 7 citations
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
2024 · 6 citations
GlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation
2023 · 3 citations
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
2023 · 3 citations
MTA-Agent: An Open Recipe for Multimodal Deep Search Agents
2026
On the Generalization Capacities of MLLMs for Spatial Intelligence
2026
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
2026
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
2025
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
2025
Scaling Agentic Reinforcement Learning For Tool-integrated Reasoning In Vlms
2025
Retrieval-augmented GUI Agents With Generative Guidelines
2025
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
2024
Top co-authors
Zeyuan Chen
· 6
Caiming Xiong
· 5
Can Qin
· 5
An Yan
· 3
Chien-Sheng Wu
· 2
Jun Wang
· 2
Le Xue
· 2
Ning Yu
· 2
Silvio Savarese
· 2
Xiangyu Peng
· 2
Xinyi Yang
· 2
Anas Awadalla
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Benchmarks
Video-Language
Embodied & Agents
Audio-Visual
Instruction Tuning
Image-Text Retrieval