Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Steven Hoi — most-cited papers & profile · Multimodal
← authors
·
overview
Steven Hoi
19
papers ·
3449
citations ·
0
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
2023 · 920 citations
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
2022 · 868 citations
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
2021 · 823 citations
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
2023 · 406 citations
Question-Guided Hybrid Convolution for Visual Question Answering
2018 · 22 citations
Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering
2018 · 17 citations
Top co-authors
Junnan Li
· 4
Dongxu Li
· 3
Caiming Xiong
· 2
Hongsheng Li
· 2
Pan Lu
· 2
Xiaogang Wang
· 2
Akhilesh Deepak Gotmare
· 1
Anthony Meng Huat Tiong
· 1
Boyang Li
· 1
Gao Peng
· 1
Haoxuan You
· 1
Junqi Zhao
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Video-Language
Image-Text Retrieval
Instruction Tuning