Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
image-text
loadingβ¦
π€
Ask AI
Awesome image-text β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
image-text
17 papers tagged image-text β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
17 papers Β· trending (default)
numbers = π₯ heat
LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
(2026)
Zhanhao Liang et al.
1.83
(1D) Ordered Tokens Enable Efficient Test-Time Search
(2026)
Zhitong Gao et al.
1.83
2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
(2025)
Wenqi Zhang et al.
1.28
Should VLMs be Pre-trained with Image Data?
(2025)
Sedrick Keh et al.
1.28
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
(2025)
Vatsal Agarwal et al.
1.28
MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
(2025)
Siyue Zhang et al.
1.28
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
(2025)
Chenhui Gou et al.
1.28
Tiny LVLM-eHub: Early Multimodal Experiments with Bard
(2023)
Wenqi Shao et al.
β
Leveraging Unpaired Data for Vision-Language Generative Models via Cycle Consistency
(2023)
Tianhong Li et al.
β
VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation
(2023)
Jinguo Zhu et al.
β
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
(2023)
Zhe Chen et al.
β
Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support
(2024)
Xiaojun Wu et al.
β
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
(2024)
Qingyun Li et al.
β
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
(2024)
Weihao Yu et al.
β
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
(2024)
Dongyang Liu et al.
β
GATE OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
(2024)
Pengfei Zhou et al.
β
Discriminative Fine-tuning of LVLMs
(2024)
Yassine Ouali et al.
β