Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Chen Li — most-cited papers & profile · Multimodal
← authors
·
overview
Chen Li
98
papers ·
412
citations ·
18
h-index
Educational Testing Service
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
2024 · 1 citations
WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization
2026
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
2026
GazeMoE: Perception of Gaze Target with Mixture-of-Experts
2026
Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress
2026
FaithSCAN: Model-Driven Single-Pass Hallucination Detection for Faithful Visual Question Answering
2026
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
2025
V-Thinker: Interactive Thinking with Images
2025
Metavla: Unified Meta Co-training For Efficient Embodied Adaption
2025
STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
2025
Arc-hunyuan-video-7b: Structured Video Comprehension Of Real-world Shorts
2025
HMR3D: Hierarchical Multimodal Representation For 3D Scene Understanding With Large Vision-language Model
2025
Know-show: Benchmarking Video-language Models On Spatio-temporal Grounded Reasoning
2025
Top co-authors
Jing Lyu
· 3
Han Zhang
· 2
Yixiao Ge
· 2
Yuying Ge
· 2
Chaodong Tong
· 1
Chenchen Zhu
· 1
Chong Sun
· 1
Enhui Wan
· 1
Guanting Dong
· 1
Honggang Zhang
· 1
Jia Xu
· 1
Jinguo Zhu
· 1
Topics
Vision-Language Models
Benchmarks
Visual QA & Reasoning
Video-Language
Embodied & Agents
Audio-Visual
Instruction Tuning
Image-Text Retrieval