Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Liang Zhao — most-cited papers & profile · Multimodal
← authors
·
overview
Liang Zhao
112
papers ·
6192
citations ·
0
h-index
Emory University · Anhui University of Science and Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
2024 · 24 citations
DreamLLM: Synergistic Multimodal Comprehension and Creation
2023 · 15 citations
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
2024 · 1 citations
STEP3-VL-10B Technical Report
2026
Describe-to-score: Text-guided Efficient Image Complexity Assessment
2025
Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs
2026
Bi-Level Prompt Optimization for Multimodal LLM-as-a-Judge
2026
Step-Audio-R1 Technical Report
2025
MiMo-VL Technical Report
2025
Multimodal Representation Learning Conditioned On Semantic Relations
2025
Cross-modal RAG: Sub-dimensional Text-to-image Retrieval-augmented Generation
2025
Perception in Reflection
2025
CLII: Visual-Text Inpainting via Cross-Modal Predictive Interaction
2024
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
2024
Top co-authors
Xiangyu Zhang
· 4
Jianjian Sun
· 3
Yuang Peng
· 3
Zheng Ge
· 3
Chengyuan Yao
· 2
Chengyue Wu
· 2
Chong Ruan
· 2
Chunrui Han
· 2
Daxin Jiang
· 2
En Yu
· 2
Haoran Wei
· 2
Haowei Zhang
· 2
Topics
Vision-Language Models
Visual QA & Reasoning
Image-Text Retrieval
Benchmarks
Video-Language
Audio-Visual
cs.SD
eess.AS
Embodied & Agents