Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Rongrong Ji — most-cited papers & profile · Multimodal
← authors
·
overview
Rongrong Ji
49
papers ·
663
citations ·
88
h-index
Xiamen University · Soochow University · Hubei University of Technology
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
2023 · 45 citations
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
2024 · 1 citations
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
2025
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
2025
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
2025
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
2025
Adapting Pre-trained Language Models to Vision-Language Tasks via Dynamic Visual Prompting
2023
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
2024
Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
2024
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
2024
Top co-authors
Chaoyou Fu
· 4
Peixian Chen
· 4
Xiaoshuai Sun
· 4
Xiawu Zheng
· 4
Mengdan Zhang
· 3
Yiyi Zhou
· 3
Xing Sun
· 2
Yunhang Shen
· 2
Yunhang Shen
· 2
Caifeng Shan
· 1
Chenyu Zhou
· 1
Erfei Cui
· 1
Topics
Vision-Language Models
Visual QA & Reasoning
Benchmarks
Video-Language
cs.CV
Image-Text Retrieval
Embodied & Agents
Instruction Tuning
Audio-Visual
cs.CL