Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Sijie Zhao — most-cited papers & profile · Multimodal
← authors
·
overview
Sijie Zhao
7
papers ·
249
citations ·
5
h-index
Tencent (China) · First Affiliated Hospital of Bengbu Medical College
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Making Llama SEE And Draw With SEED Tokenizer
2023 · 203 citations
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
2023 · 6 citations
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
2024 · 3 citations
OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation
2026
VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation
2023
Topics
Vision-Language
Model Architecture
Fine-Tuning
image
performance
In-Context Learning
Multi-Agent
Tool Use
Diffusion Models
Text-to-Video