← all papers · overview

Specattn: Co-designing Sparse Attention With Self-speculative Decoding

Abstract

Long-context large language model (LLM) inference has become the norm for today's AI applications. However, it is severely bottlenecked by the increasing memory demands of its KV cache. Previous works have shown that self-speculative decoding with sparse attention, where tokens are drafted using a subset of the KV cache and verified in parallel with full KV cache, speeds up inference in a lossless

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).