Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jiajun Wu — most-cited papers & profile · Multimodal
← authors
·
overview
Jiajun Wu
136
papers ·
3942
citations ·
58
h-index
Beijing University of Posts and Telecommunications · Zhejiang University of Science and Technology · Shantou University · ZheJiang Academy of Agricultural Sciences · Massachusetts Institute of Technology · Stanford University
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Retrospectives on the Embodied AI Workshop
2022 · 19 citations
HourVideo: 1-Hour Video-Language Understanding
2024 · 5 citations
Learning Situated Awareness in the Real World
2026
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models
2026
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
2026
Thinking with Spatial Code for Physical-World Video Reasoning
2026
TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
2025
ENACT: Evaluating Embodied Cognition With World Modeling Of Egocentric Interaction
2025
Visually Descriptive Language Model for Vector Graphics Reasoning
2024
Top co-authors
Manling Li
· 3
Joy Hsu
· 2
Li Fei-Fei
· 2
Alan Yuille
· 1
Alexander Toshev
· 1
Ali Farhadi
· 1
Aniruddha Kembhavi
· 1
Chengshu Li
· 1
Chuang Gan
· 1
Danyang Zhang
· 1
Devendra Singh Chaplot
· 1
Eric Hanchen Jiang
· 1
Topics
Visual QA & Reasoning
Vision-Language Models
Benchmarks
Video-Language
Embodied & Agents