← all papers · overview

Don't Waste Bits! Adaptive Kv-cache Quantization For Lightweight On-device Llms

Abstract

Large Language Models (LLMs) have achieved remarkable progress across reasoning, generation, and decision-making tasks, yet deploying them on mobile, embedded, and edge devices remains particularly challenging. On-device LLM inference is heavily constrained by the memory and bandwidth overhead of the key-value (KV) cache, which grows linearly with context length and often dominates decoding cost.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).