← all papers · overview

ARKV: Adaptive And Resource-efficient KV Cache Management Under Limited Memory Budget For Long-context Inference In Llms

Abstract

Large Language Models (LLMs) are increasingly deployed in scenarios demanding ultra-long context reasoning, such as agentic workflows and deep research understanding. However, long-context inference is constrained by the KV cache, a transient memory structure that grows linearly with sequence length and batch size, quickly dominating GPU memory usage. Existing memory reduction techniques, includin

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).