Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Zhen Yang — most-cited papers & profile · Multimodal
← authors
·
overview
Zhen Yang
89
papers ·
756
citations ·
37
h-index
Qingdao University · Nanjing Normal University · Gansu Provincial Maternal and Child Health Hospital · Henan Normal University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
ViLTA: Enhancing Vision-Language Pre-training through Textual Augmentation
2023 · 8 citations
CLIP$^2$: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data
2023 · 3 citations
Kimi K2.5: Visual Agentic Intelligence
2026 · 1 citations
MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?
2025
Understanding Temporal Logic Consistency In Video-language Models Through Cross-modal Attention Discriminability
2025
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
2025
Mmgeolm: Hard Negative Contrastive Learning For Fine-grained Geometric Understanding In Large Multimodal Models
2025
Ui2code^n: A Visual Language Model For Test-time Scalable Interactive Ui-to-code Generation
2025
Attribute Guidance With Inherent Pseudo-label For Occluded Person Re-identification
2025
Top co-authors
Yutao Zhang
· 3
Angang Du
· 2
Bin Xu
· 2
Bohong Yin
· 2
Bowei Xing
· 2
Bowen Qu
· 2
Chao Hong
· 2
Cheng Li
· 2
Chenzhuang Du
· 2
Chuning Tang
· 2
Chu Wei
· 2
Dan Ye
· 2
Topics
Vision-Language Models
Benchmarks
Video-Language
Visual QA & Reasoning
Embodied & Agents
Image-Text Retrieval