Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Zihan Xu — most-cited papers & profile · Multimodal
← authors
·
overview
Zihan Xu
12
papers ·
3
citations ·
8
h-index
Dalian University of Technology · Nanjing University of Aeronautics and Astronautics
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
VT-LVLM-AR: A Video-temporal Large Vision-language Model Adapter For Fine-grained Action Recognition In Long-term Videos
2025
Vl-medguide: A Visual-linguistic Large Model For Intelligent And Explainable Skin Disease Auxiliary Diagnosis
2025
Top co-authors
Jialei Xie
· 1
Kaining Li
· 1
Kexin Yu
· 1
Shuwei He
· 1
Topics
Vision-Language Models
Video-Language
Visual QA & Reasoning