Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
compression
loadingβ¦
π€
Ask AI
Awesome compression β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
compression
24 papers tagged compression β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
24 papers Β· trending (default)
numbers = π₯ heat
On-Policy Self-Distillation for Reasoning Compression
(2026)
Hejian Sang et al.
1.94
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
(2026)
Ziyun Zeng et al.
1.94
LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning
(2026)
Mengmeng Ji et al.
1.94
EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management
(2026)
Zherui Yang et al.
1.94
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
(2026)
Guozhen Zhang et al.
1.94
RoPE-Aware Bit Allocation for KV-Cache Quantization
(2026)
Fengfeng Liang et al.
1.94
Small Vision-Language Models are Smart Compressors for Long Video Understanding
(2026)
Junjie Fei et al.
1.83
A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
(2026)
Jincheng Ren et al.
1.83
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
(2026)
Aditya Kumar Singh et al.
1.72
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
(2025)
Yuri Kuratov et al.
1.28
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
(2025)
Weiming Ren et al.
1.28
QwenLong-CPRS: Towards infty-LLMs with Dynamic Context Optimization
(2025)
Weizhou Shen et al.
1.28
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
(2025)
Senqiao Yang et al.
1.28
Pruning the Unsurprising: Efficient Code Reasoning via First-Token Surprisal
(2025)
Wenhao Zeng et al.
1.28
Retrieval-augmented reasoning with lean language models
(2025)
Ryan Sze-Yin Chan et al.
1.28
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
(2025)
Shufan Li et al.
1.28
Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
(2025)
Tianfan Peng et al.
1.28
Vcc: Scaling Transformers to 128K Tokens or More by Prioritizing Important Tokens
(2023)
Zhanpeng Zeng et al.
β
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
(2024)
Yi-Fan Zhang et al.
β
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
(2024)
Can Qin et al.
β
A Web-Based Solution for Federated Learning with LLM-Based Automation
(2024)
Chamith Mawela et al.
β
A Survey of Small Language Models
(2024)
Chien Van Nguyen et al.
β
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems
(2024)
Jiejun Tan et al.
β
Whisper-GPT: A Hybrid Representation Audio Large Language Model
(2024)
Prateek Verma
β