Awesome Large Language Models
π
Papers
π§
Topics
π₯
Trending
πΊοΈ
Map
π
Leaderboards
π
Learn
π€
Ask AI
β―
More
π₯
Authors
π
Reading Packs
π
Datasets
π οΈ
Tools
π°
News
π
Blogs
βοΈ
Newsletter
π―
Research Radar
π
Saved
+ Add Paper
βΎ
β
β all topics
overview
cache
loadingβ¦
π€
Ask AI
Awesome cache β curated papers, datasets & benchmarks Β· Awesome Large Language Models
β all topics
overview
cache
22 papers tagged cache β re-sort below
Papers
π₯ Trending (default)
π Most cited
π Newest first
π€ A β Z by title
22 papers Β· trending (default)
numbers = π₯ heat
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
(2025)
Minsoo Kim et al.
5.94
OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion Models
(2026)
Zhaoyuan He et al.
4.33
ForgettingOT: Certified Speculative Batching from Sinkhorn's Projective Forgetting
(2026)
Xinyang Wen
4.33
When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers
(2026)
Yushi Sun et al.
1.94
WaveFilter: Enhancing the Long-Context Capability of Diffusion LLMs via Wavelet-Guided KV Cache Filtering
(2026)
Jinnan Yang et al.
1.89
Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching
(2026)
Qianli Ma et al.
1.89
Wan-Streamer v0.2: Higher Resolution, Same Latency
(2026)
Lianghua Huang et al.
1.89
KVpop -- Key-Value Cache Compression with Predictive Online Pruning
(2026)
Lukas Hauzenberger et al.
1.89
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
(2026)
Hongyu Liu et al.
1.89
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing
(2026)
Yongqi An et al.
1.83
LearnedCache: eBPF-Integrated Perceptron-Based Eviction Policies for the Linux Page Cache
(2026)
Zejia Qi
1.83
GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs
(2026)
Junjie Peng et al.
1.83
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
(2026)
Chuangtao Chen et al.
1.78
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
(2026)
Yingsheng Geng et al.
1.72
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
(2026)
Yichun Xu et al.
1.72
Joint Encoding of KV-Cache Blocks for Scalable LLM Serving
(2026)
Joseph Kampeas and Emir Haleva
1.61
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
(2026)
Zihan Wang et al.
1.61
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
(2026)
Jang-Hyun Kim et al.
1.61
Randomization Boosts KV Caching, Learning Balances Query Load: A Joint Perspective
(2026)
Fangzhou Wu et al.
1.61
KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
(2025)
Aomufei Yuan et al.
1.56
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
(2025)
Bo Jiang et al.
1.56
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
(2025)
Ayan Sengupta et al.
1.39