Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Boyang Li — most-cited papers & profile · Multimodal
← authors
·
overview
Boyang Li
13
papers ·
440
citations ·
32
h-index
Hebei Medical University · Tianjin University · Peking University · Directorate for Cultural Heritage · Wuhan University · Robotics Research (United States) · Renmin University of China · Shaanxi Normal University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
2023 · 406 citations
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training
2022 · 6 citations
CAT Merging: A Training-free Approach For Resolving Conflicts In Model Merging
2025
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
2023
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
2024
Top co-authors
Anthony Meng Huat Tiong
· 2
Jiayun Luo
· 2
Junnan Li
· 2
Leonid Sigal
· 2
Dongxu Li
· 1
Junqi Zhao
· 1
Mir Rayat Imtiaz Hossain
· 1
Pascale Fung
· 1
Pritam Sarkar
· 1
Qingyong Li
· 1
Siddhesh Khandelwal
· 1
Silvio Savarese
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Video-Language
Benchmarks
Instruction Tuning
Audio-Visual