← all papers · overview

Locret: Enhancing Eviction In Long-context LLM Inference With Trained Retaining Heads On Consumer-grade Devices

Abstract

Scaling the input context length of a large language model (LLM) incurs a significant increase in computation cost and memory footprint to maintain the attention key-value (KV) cache. Existing KV cache compression methods suffer from inefficient compression strategies and limited memory reduction effects, making it difficult for LLMs to conduct long-context inference on consumer-grade devices, esp

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).