Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Bo Wang — most-cited papers & profile · Multimodal
← authors
·
overview
Bo Wang
121
papers ·
1249
citations ·
19
h-index
Henan Institute of Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Knowledge Condensation and Reasoning for Knowledge-based VQA
2024 · 2 citations
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
2024 · 2 citations
Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models
2026
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
2026
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
2026
A Survey on Efficient Vision-Language-Action Models
2025
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
2025
How Far Are We From Generating Missing Modalities With Foundation Models?
2025
Crosshoi-bench: A Unified Benchmark For HOI Evaluation Across Vision-language Models And Hoi-specific Methods
2025
Hola: Zero-shot HOI Detection With Low-rank Decomposed VLM Feature Adaptation
2025
CAS-IQA: Teaching Vision-language Models For Synthetic Angiography Quality Assessment
2025
Enhanced Continual Learning of Vision-Language Models with Model Fusion
2025
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
2024
AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis
2024
Top co-authors
Haoyang Huang
· 2
Linghe Kong
· 2
Nan Duan
· 2
Robby T. Tan
· 2
Te Yang
· 2
Weiran Huang
· 2
Wenbo Li
· 2
Yan Li
· 2
Bin Li
· 1
Bohan Zeng
· 1
Dongze Hao
· 1
Guohui Zhang
· 1
Topics
Vision-Language Models
Benchmarks
Visual QA & Reasoning
Video-Language
Embodied & Agents
Instruction Tuning
Image-Text Retrieval