← all papers · overview

Training-free Exponential Context Extension Via Cascading KV Cache

Abstract

The transformer's context window is vital for tasks such as few-shot learning and conditional generation as it preserves previous tokens for active memory. However, as the context lengths increase, the computational costs grow quadratically, hindering the deployment of large language models (LLMs) in real-world, long sequence scenarios. Although some recent key-value caching (KV Cache) methods off

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).