Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Zhengyuan Yang — most-cited papers & profile · Multimodal
← authors
·
overview
Zhengyuan Yang
26
papers ·
1140
citations ·
29
h-index
National University of Defense Technology · Stomatology Hospital · Microsoft Research (United Kingdom)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
2021 · 46 citations
Multimodal Foundation Models: From Specialists to General-Purpose Assistants
2023 · 24 citations
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
2025 · 1 citations
Equivariant Similarity for Vision-Language Foundation Models
2023 · 1 citations
Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents
2025
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
2025
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
2025
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
2024
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
2024
Top co-authors
Linjie Li
· 6
Lijuan Wang
· 4
Chung-Ching Lin
· 3
Kevin Lin
· 3
Furong Huang
· 2
Mike Zheng Shou
· 2
Xiyao Wang
· 2
Zhe Gan
· 2
Zicheng Liu
· 2
Alex Jinpeng Wang
· 1
Alex Jinpeng Wang
· 1
Chao Feng
· 1
Topics
Vision-Language Models
Benchmarks
cs.CV
Visual QA & Reasoning
Image-Text Retrieval
Video-Language
cs.CL
cs.LG