Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jing Li — most-cited papers & profile · Multimodal
← authors
·
overview
Jing Li
107
papers ·
1909
citations ·
36
h-index
Harbin Institute of Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Multimodal Transformer with Multi-View Visual Representation for Image Captioning
2019 · 30 citations
VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment
2022 · 11 citations
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model
2026
PPU-Bench:Real World Benchmark for Personalized Partial Unlearning in Vision Language Models
2026
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
2026
PRISM of Opinions: A Persona-Reasoned Multimodal Framework for User-centric Conversational Stance Detection
2025
Mdaif: Robust One-stop Multi-degradation-aware Image Fusion With Language-driven Semantics
2025
TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
2025
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
2025
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
2025
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
2025
Multimodal Reasoning with Multimodal Knowledge Graph
2024
Top co-authors
Wenya Wang
· 2
Yu Yin
· 2
Zhe Hu
· 2
Bingbing Wang
· 1
Bing Zhao
· 1
Chunzhao Xie
· 1
Cuiyun Gao
· 1
Guodong DU
· 1
Han Wu
· 1
Hao Zhang
· 1
Hardik Shah
· 1
Hou Pong Chan
· 1
Topics
Vision-Language Models
Benchmarks
Visual QA & Reasoning
Video-Language
Image-Text Retrieval