Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Jiuxiang Gu — most-cited papers & profile · Multimodal
← authors
·
overview
Jiuxiang Gu
18
papers ·
117
citations ·
27
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models
2017 · 37 citations
Delving into Out-of-Distribution Detection with Vision-Language Representations
2022 · 31 citations
SelfDoc: Self-Supervised Document Representation Learning
2021 · 7 citations
TRINS: Towards Multimodal Language Models that Can Read
2024 · 1 citations
MENTOR: Efficient Multimodal-conditioned Tuning For Autoregressive Vision Generation Models
2025
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
2025
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
2025
Towards Visual Text Grounding of Multimodal Large Language Model
2025
Learning the Visualness of Text Using Large Vision-Language Models
2023
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
2024
MMR: Evaluating Reading Ability of Large Multimodal Models
2024
Top co-authors
Jian Chen
· 4
Ruiyi Zhang
· 4
Changyou Chen
· 3
Tong Sun
· 3
Handong Zhao
· 2
Jason Kuen
· 2
Yufan Zhou
· 2
Yufan Zhou
· 2
Aditya Grover
· 1
and Yixuan Li
· 1
Ani Nenkova
· 1
Chenguang Wang
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Benchmarks
Instruction Tuning
Video-Language
Audio-Visual
Image-Text Retrieval