Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
document
loadingβ¦
π€
Ask AI
Awesome document β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
document
24 papers tagged document β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
24 papers Β· trending (default)
numbers = π₯ heat
More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
(2025)
Shahar Levy et al.
2.87
NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval
(2026)
Zhuchenyang Liu et al.
1.94
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents
(2026)
Udi Barzelay et al.
1.94
Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG
(2026)
Yiqun Sun et al.
1.94
Counting as a minimal probe of language model reliability
(2026)
Tianxiang Dai et al.
1.94
Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing
(2026)
Nilesh Jain
1.94
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
(2026)
Peiyang Liu et al.
1.89
PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents
(2026)
Bihui Yu et al.
1.89
DocAtlas: Multilingual Document Understanding Across 80+ Languages
(2026)
Ahmed Heakl et al.
1.89
FireRed-OCR Technical Report
(2026)
Hao Wu et al.
1.78
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
(2026)
Daxiang Dong et al.
1.78
MMDocIR: Benchmarking Multi-Modal Retrieval for Long Documents
(2025)
Kuicai Dong et al.
1.28
PlainQAFact: Automatic Factuality Evaluation Metric for Biomedical Plain Language Summaries Generation
(2025)
Zhiwen You et al.
1.28
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
(2025)
Ahmed Nassar et al.
1.28
Distillation and Refinement of Reasoning in Small Language Models for Document Re-ranking
(2025)
Chris Samarinas et al.
1.28
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
(2025)
Wenxuan Shen et al.
1.28
Retrieval-augmented reasoning with lean language models
(2025)
Ryan Sze-Yin Chan et al.
1.28
Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
(2025)
Yurun Chen et al.
1.28
Document Understanding, Measurement, and Manipulation Using Category Theory
(2025)
Jared Claypoole et al.
1.28
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
(2025)
Shengyuan Ding et al.
1.28
Needle In A Multimodal Haystack
(2024)
Weiyun Wang et al.
β
L-CiteEval: Do Long-Context Models Truly Leverage Context for Responding?
(2024)
Zecheng Tang et al.
β
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
(2024)
Sara Ghaboura et al.
β
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding
(2024)
Jaemin Cho et al.
β