← all papers · overview

Near-oracle KV Selection Via Pre-hoc Sparsity For Long-context Inference

Abstract

A core bottleneck in large language model (LLM) inference is the cost of attending over the ever-growing key-value (KV) cache. Although near-oracle top-k KV selection can preserve the quality of dense attention while sharply reducing computation and bandwidth, existing sparse methods generally rely on posterior heuristics, i.e., selectors conditioned on observed attention or proxy scores. Such con

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).