Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Ji-Rong Wen — most-cited papers & profile · Multimodal
← authors
·
overview
Ji-Rong Wen
64
papers ·
2032
citations ·
76
h-index
Beijing Academy of Artificial Intelligence · Next Generation Technology (United States) · Renmin University of China · Chinese Academy of Governance
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Evaluating Object Hallucination in Large Vision-Language Models
2023 · 362 citations
Recursive Visual Attention in Visual Dialog
2018 · 5 citations
Learning to Answer Questions in Dynamic Audio-Visual Scenarios
2022 · 5 citations
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
2023 · 4 citations
COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval
2022 · 2 citations
Zero-shot Visual Question Answering with Language Model Feedback
2023 · 1 citations
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
2025
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
2025
Progressive Multimodal Reasoning via Active Retrieval
2024
Top co-authors
Wayne Xin Zhao
· 4
Yifan Du
· 3
Zhiwu Lu
· 2
Chenghao Zhang
· 1
Chenliang Xu
· 1
Chuyuan Wang
· 1
Dawei Gao
· 1
Di Hu
· 1
Guangyao Li
· 1
Guanting Dong
· 1
Guanzhong Wang
· 1
Hangyu Guo
· 1
Topics
Visual QA & Reasoning
Vision-Language Models
cs.CL
cs.CV
Benchmarks
Image-Text Retrieval
Video-Language
cs.AI
Audio-Visual
cs.MM