Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Haoyang Huang — most-cited papers & profile · Multimodal
← authors
·
overview
Haoyang Huang
21
papers ·
200
citations ·
15
h-index
Qinghai University · Shanghai Jiao Tong University · Nanjing University of Science and Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
2020 · 169 citations
XGPT: Cross-modal Generative Pre-Training for Image Captioning
2020 · 20 citations
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
2026
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
2026
Top co-authors
Lei Ji
· 2
Ming Zhou
· 2
Nan Duan
· 2
Taroon Bharti
· 2
Bohan Zeng
· 1
Botian Shi
· 1
Bo Wang
· 1
Bo Wang
· 1
Dongdong Zhang
· 1
Edward Cui
· 1
Guohui Zhang
· 1
Guoqing Huang
· 1
Topics
Benchmarks
Vision-Language Models
Visual QA & Reasoning
Embodied & Agents
Instruction Tuning
Video-Language
Image-Text Retrieval