← all papers · overview

Tokenselect: Efficient Long-context Inference And Length Extrapolation For Llms Via Dynamic Token-level KV Cache Selection

Abstract

Rapid advances in Large Language Models (LLMs) have spurred demand for processing extended context sequences in contemporary applications. However, this progress faces two challenges: performance degradation due to sequence lengths out-of-distribution, and excessively long inference times caused by the quadratic computational complexity of attention. These issues limit LLMs in long-context scenari

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).