Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Caiming Xiong — most-cited papers & profile · Multimodal
← authors
·
overview
Caiming Xiong
105
papers ·
4726
citations ·
69
h-index
Salesforce (United States)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
2022 · 868 citations
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
2021 · 823 citations
GlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation
2023 · 3 citations
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
2023 · 3 citations
LATTE: Learning to Think with Vision Specialists
2024 · 1 citations
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
2025
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
2024
Top co-authors
Junnan Li
· 3
Silvio Savarese
· 3
Can Qin
· 2
Dongxu Li
· 2
Juan Carlos Niebles
· 2
Manli Shu
· 2
Shafiq Joty
· 2
Steven Hoi
· 2
Zeyuan Chen
· 2
Akhilesh Deepak Gotmare
· 1
Anas Awadalla
· 1
An Yan
· 1
Topics
Vision-Language Models
Benchmarks
Visual QA & Reasoning
Image-Text Retrieval
Video-Language
Audio-Visual
Instruction Tuning