Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Gangyan Zeng — most-cited papers & profile · Multimodal
← authors
·
overview
Gangyan Zeng
7
papers ·
4
citations
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
2024 · 1 citations
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
2026
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
2025
Top co-authors
Can Ma
· 1
Can Ma
· 1
Chengquan Zhang
· 1
Daiqing Wu
· 1
Gengluo Li
· 1
Hangui Lin
· 1
Han Hu
· 1
Harry Yang
· 1
Huawen Shen
· 1
Huawen Shen
· 1
Nicu Sebe
· 1
Pengyuan Lyu
· 1
Topics
Visual QA & Reasoning
Image-Text Retrieval
Benchmarks
Vision-Language Models
Video-Language