Awesome Multimodal
📄
Papers
🧭
Topics
🔥
Trending
🗺️
Map
🏆
Leaderboards
🎓
Learn
🤖
Ask AI
⋯
More
👥
Authors
📚
Reading Packs
📊
Datasets
🛠️
Tools
📰
News
📝
Blogs
✉️
Newsletter
🎯
Research Radar
🔖
Saved
+ Add Paper
☾
☀
← authors
·
overview
Loading author…
🤖
Ask AI
Xiaoyu Shen — most-cited papers & profile · Multimodal
← authors
·
overview
Xiaoyu Shen
35
papers ·
140
citations ·
0
h-index
Google Scholar ↗
Semantic Scholar ↗
OpenAlex ↗
Most-cited papers
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
2026
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
2026
Beyond Global Similarity: Towards Fine-Grained, Multi-Condition Multimodal Retrieval
2026
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
2026
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
2026
UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
2026
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
2026
MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
2025
$\mathcalVisi\mathcalPruner$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
2025
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
2025
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
2025
Top co-authors
Junlong Tong
· 6
Hao Wu
· 5
Yingqi Fan
· 4
Yunpu Ma
· 4
Anhao Zhao
· 3
Jinlan Fu
· 3
Junyan Lin
· 3
Hao Fei
· 2
See-Kiong Ng
· 2
Xipeng Qiu
· 2
XuDong Wang
· 2
Bryan Hooi
· 1
Topics
Vision-Language Models
Benchmarks
Video-Language
Visual QA & Reasoning
Audio-Visual
Image-Text Retrieval