← all papers · overview

BUZZ: Beehive-structured Sparse KV Cache With Segmented Heavy Hitters For Efficient LLM Inference

Abstract

Large language models (LLMs) are essential in natural language processing but often struggle with inference speed and computational efficiency, limiting real-time deployment. The key-value (KV) cache mechanism reduces computational overhead in transformer models, but challenges in maintaining contextual understanding remain. In this paper, we propose BUZZ, a novel KV caching algorithm that leverag

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).