Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Junnan Li — most-cited papers & profile · Multimodal
← authors
·
overview
Junnan Li
9
papers ·
3062
citations ·
19
h-index
University of Science and Technology Liaoning · Huaihua University · University of Electronic Science and Technology of China · Runze (China) · Zhejiang Industry Polytechnic College · University of Rochester
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
2023 · 920 citations
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
2022 · 868 citations
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
2021 · 823 citations
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
2023 · 406 citations
Align and Prompt: Video-and-Language Pre-training with Entity Prompts
2021 · 17 citations
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training
2022 · 6 citations
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
2023 · 3 citations
Top co-authors
Dongxu Li
· 5
Steven Hoi
· 4
Caiming Xiong
· 3
Silvio Savarese
· 3
Anthony Meng Huat Tiong
· 2
Boyang Li
· 2
Juan Carlos Niebles
· 2
Shafiq Joty
· 2
Steven C.H. Hoi
· 2
Akhilesh Deepak Gotmare
· 1
Artemis Panagopoulou
· 1
Hongdong Li
· 1
Topics
Vision-Language Models
Video-Language
Visual QA & Reasoning
Image-Text Retrieval
Instruction Tuning
Audio-Visual
Benchmarks