Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Wanli Ouyang — most-cited papers & profile · Multimodal
← authors
·
overview
Wanli Ouyang
39
papers ·
239
citations ·
101
h-index
Chinese University of Hong Kong · Beijing Academy of Artificial Intelligence · Shanghai Artificial Intelligence Laboratory
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Dense Video Captioning Using Graph-based Sentence Summarization
2025 · 48 citations
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
2023 · 40 citations
Show, Tell And Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
2025 · 36 citations
Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models
2023 · 6 citations
Dense Connector for MLLMs
2024 · 1 citations
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
2024 · 1 citations
Interleaving Reasoning for Better Text-to-Image Generation
2025
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
2025
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?
2025
Top co-authors
Dong Xu
· 2
Jingdong Wang
· 2
Wenhao Wu
· 2
Wenxuan Huang
· 2
Zhiwang Zhang
· 2
Bangyan Li
· 1
Chuanqi Tan
· 1
Dian Zheng
· 1
Dingning Liu
· 1
Di Zhang
· 1
Dongzhan Zhou
· 1
Fan Zhang
· 1
Topics
Vision-Language Models
Video-Language
Benchmarks
Visual QA & Reasoning
Instruction Tuning
Audio-Visual
Image-Text Retrieval