← all papers · overview

Zipage: Maintain High Request Concurrency For LLM Reasoning Through Compressed Pagedattention

Abstract

With reasoning becoming the generative paradigm for large language models (LLMs), the memory bottleneck caused by KV cache during the decoding phase has become a critical factor limiting high-concurrency service. Although existing KV cache eviction methods address the memory issue, most of them are impractical for industrial-grade applications. This paper introduces Compressed PagedAttention, a me

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).