Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jianfei Cai — most-cited papers & profile · Multimodal
← authors
·
overview
Jianfei Cai
39
papers ·
380
citations ·
0
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models
2017 · 37 citations
Causal Attention for Vision-Language Tasks
2021 · 5 citations
Auto-Parsing Network for Image Captioning and Visual Question Answering
2021 · 1 citations
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
2024
How Well Can Vision Language Models See Image Details?
2024
Top co-authors
Hanwang Zhang
· 2
Xu Yang
· 2
Chongyang Gao
· 1
Deyao Zhu
· 1
Gang Wang
· 1
Jiuxiang Gu
· 1
Li Niu
· 1
Mohamed Elhoseiny
· 1
Shafiq Joty
· 1
Ziyu Ma
· 1
Topics
Vision-Language Models
Audio-Visual
Visual QA & Reasoning
Video-Language
Image-Text Retrieval
Benchmarks