Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
cross-modal
loadingβ¦
π€
Ask AI
Awesome cross-modal β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
cross-modal
21 papers tagged cross-modal β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
21 papers Β· trending (default)
numbers = π₯ heat
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
(2025)
Jeonghyeon Kim et al.
4.47
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
(2026)
Hengyu Shen et al.
1.94
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
(2026)
Ruiheng Zhang et al.
1.94
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
(2026)
Dan Ben-Ami et al.
1.94
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
(2026)
Feng Han et al.
1.94
GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods
(2026)
Sujay Belsare et al.
1.94
Small Vision-Language Models are Smart Compressors for Long Video Understanding
(2026)
Junjie Fei et al.
1.83
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
(2025)
Mohammad Mahdi Abootorabi et al.
1.28
MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching
(2025)
Fabian David Schmidt et al.
1.28
Multimodal Language Modeling for High-Accuracy Single Cell Transcriptomics Analysis and Generation
(2025)
Yaorui Shi et al.
1.28
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
(2025)
Qi Qin et al.
1.28
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
(2025)
Xinjie Zhang et al.
1.28
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
(2025)
Lin Long et al.
1.28
RAG-Anything: All-in-One RAG Framework
(2025)
Zirui Guo et al.
1.28
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
(2025)
Siyin Wang et al.
1.28
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
(2025)
Yongyuan Liang et al.
1.28
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks
(2023)
Haiyang Xu et al.
β
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
(2024)
Haoji Zhang et al.
β
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
(2024)
Hang Hua et al.
β
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
(2024)
Sicong Leng et al.
β
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
(2024)
Jaemin Cho et al.
β