Awesome Multimodal
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β authors
Β·
overview
Loading authorβ¦
π€
Ask AI
Yuanhan Zhang β most-cited papers & profile Β· Multimodal
β authors
Β·
overview
Yuanhan Zhang
7
papers Β·
34
citations
Google Scholar β
Semantic Scholar β
OpenAlex β
Most-cited papers
LLaVA-OneVision: Easy Visual Task Transfer
2024 Β· 31 citations
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
2023 Β· 3 citations
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
2025
Otter: A Multi-Modal Model with In-Context Instruction Tuning
2023
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
2023
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
2024
LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
2024
Topics
Evaluation
Fine-Tuning
Model Architecture
Vision-Language
Training Techniques
In-Context Learning
Perception
Manipulation
Human-Robot Interaction
3D Vision