Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Yan Shu — most-cited papers & profile · Multimodal
← authors
·
overview
Yan Shu
9
papers ·
0
citations ·
0
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
2025
Timescope: Towards Task-oriented Temporal Grounding In Long Videos
2025
Task-aware KV Compression For Cost-effective Long Video Understanding
2025
Earthmind: Leveraging Cross-sensor Data For Advanced Earth Observation Interpretation With A Unified Multimodal LLM
2025
Vidtext: Towards Comprehensive Evaluation For Video Text Understanding
2025
Top co-authors
Minghao Qin
· 2
Bin Ren
· 1
Gangyan Zeng
· 1
Hangui Lin
· 1
Harry Yang
· 1
Jing Wang
· 1
Nicu Sebe
· 1
Peitian Zhang
· 1
Ser-Nam Lim
· 1
Xiangrui Liu
· 1
Yan Li
· 1
Yan Zhang
· 1
Topics
Benchmarks
Visual QA & Reasoning
Video-Language
Vision-Language Models
Image-Text Retrieval