Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Wayne Xin Zhao — most-cited papers & profile · Multimodal
← authors
·
overview
Wayne Xin Zhao
71
papers ·
2375
citations ·
56
h-index
Renmin University of China
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Evaluating Object Hallucination in Large Vision-Language Models
2023 · 362 citations
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
2021 · 85 citations
A Survey of Vision-Language Pre-Trained Models
2022 · 43 citations
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
2023 · 4 citations
Zero-shot Visual Question Answering with Language Model Feedback
2023 · 1 citations
Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
2025
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
2025
Top co-authors
Ji-Rong Wen
· 4
Yifan Du
· 4
Junyi Li
· 2
Ruihua Song
· 2
and Ji-Rong Wen
· 1
Anwen Hu
· 1
Baogui Xu
· 1
Chuhao Jin
· 1
Chuyuan Wang
· 1
Danyang Hou
· 1
Dawei Gao
· 1
Guangzhen Liu
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Video-Language
Benchmarks
Instruction Tuning
Image-Text Retrieval
Audio-Visual