← all papers · overview

RAP: Kv-cache Compression Via Rope-aligned Pruning

Abstract

Long-context inference in large language models is increasingly bottlenecked by the memory and compute cost of the KV-Cache. Low-rank factorization compresses KV projections by writing , where A produces latent KV states and B can be absorbed into downstream weights. In modern RoPE-based LLMs, this absorption fails: RoPE forces latent KV states to be reconstructed to full dimens

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).