← all papers · overview

Structkv: Preserving The Structural Skeleton For Scalable Long-context Inference

Abstract

As Large Language Models (LLMs) scale to support context windows exceeding one million tokens, the linear growth of Key-Value (KV) cache imposes severe memory capacity and bandwidth bottlenecks, constraining the efficiency of long-context inference. Existing compression approaches typically prioritize tokens based on local saliency metrics to decouple prefill computation from decoding memory. Howe

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).