Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Shuohuan Wang — most-cited papers & profile · Multimodal
← authors
·
overview
Shuohuan Wang
11
papers ·
46
citations ·
2
h-index
Baidu (China)
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
ERNIE-UniX2: A Unified Cross-lingual Cross-modal Framework for Understanding and Generation
2022 · 3 citations
ERNIE 5.0 Technical Report
2026 · 2 citations
Autoregressive Pre-Training on Pixels and Texts
2024 · 1 citations
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
2026
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
2026
V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
2025
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding
2025
Top co-authors
Hua Wu
· 6
Yu Sun
· 6
Haifeng Wang
· 5
Zhenyu Zhang
· 4
Naibin Gu
· 3
Linhao Yu
· 2
Weichong Yin
· 2
Yao Chen
· 2
Yekun Chai
· 2
Bingjin Chen
· 1
Bin Li
· 1
Bin Shan
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Benchmarks
Audio-Visual
Video-Language
Image-Text Retrieval