← all papers · overview

Longspec: Long-context Lossless Speculative Decoding With Efficient Drafting And Verification

Abstract

As Large Language Models (LLMs) can now process extremely long contexts, efficient inference over these extended inputs has become increasingly important, especially for emerging applications like LLM agents that highly depend on this capability. Speculative decoding (SD) offers a promising lossless acceleration technique compared to lossy alternatives such as quantization and model cascades. Howe

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).