← all papers · overview

Quasar: Quantized Self-speculative Acceleration For Rapid Inference Via Memory-efficient Verification

Abstract

Speculative Decoding (SD) has emerged as a premier technique for accelerating Large Language Model (LLM) inference by decoupling token generation into rapid drafting and parallel verification. While recent advancements in self-speculation and lookahead decoding have successfully minimized drafting overhead, they have shifted the primary performance bottleneck to the verification phase. Since verif

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).