Awesome Large Language Models
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Zhe Gan — most-cited papers & profile · Large Language Models
← authors
·
overview
Zhe Gan
37
papers ·
1281
citations ·
59
h-index
Gannan Medical University · Peking University Shenzhen Hospital · First Affiliated Hospital of Gannan Medical University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
2024 · 29 citations
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
2025
DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
2025
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
2025
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
2025
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
2024
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
2024
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
2024
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
2024
MM-Ego: Towards Building Egocentric Multimodal LLMs
2024
Top co-authors
Yinfei Yang
· 10
Haotian Zhang
· 7
Bowen Zhang
· 5
Peter Grasch
· 4
Haoxuan You
· 3
Mingfei Gao
· 3
Philipp Dufter
· 3
Wenze Hu
· 3
Xianzhi Du
· 3
Yanghao Li
· 3
Zirui Wang
· 3
Afshin Dehghan
· 2
Topics
Vision-Language
Training Techniques
Fine-Tuning
Evaluation
Model Architecture
Code
Efficiency
RAG
Prompting
QA