Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yang Zhang — most-cited papers & profile · Multimodal
← authors
·
overview
Yang Zhang
242
papers ·
8045
citations ·
46
h-index
Nanjing University of Aeronautics and Astronautics
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Visual hallucination detection in large vision-language models via evidential conflict
2025 · 2 citations
Rethinking The Text-vision Reasoning Imbalance In Mllms Through The Lens Of Training Recipes
2025
Reconstruction as a Bridge for Event-Based Visual Question Answering
2025
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
2025
Bridging The Gap In Vision Language Models In Identifying Unsafe Concepts Across Modalities
2025
Visualtrap: A Stealthy Backdoor Attack On GUI Agents Via Visual Grounding Manipulation
2025
Seeing The Unseen: Towards Zero-shot Inspection For Wind Turbine Blades Using Knowledge-augmented Vision Language Models
2025
Top co-authors
Boxin Shi
· 1
Boyu Li
· 1
Farhad Imani
· 1
Guangnan Ye
· 1
Guanyu Yao
· 1
Hanyue Lou
· 1
Jian Wu
· 1
Jiayi Zhou
· 1
Jindong Gu
· 1
Liping Jing
· 1
Michael Backes
· 1
Philip Torr
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Benchmarks
Video-Language
Audio-Visual
Embodied & Agents
Image-Text Retrieval