← all papers · overview

Speeding Up Speculative Decoding Via Sequential Approximate Verification

Abstract

Speculative Decoding (SD) is a recently proposed technique for faster inference using Large Language Models (LLMs). SD operates by using a smaller draft LLM for autoregressively generating a sequence of tokens and a larger target LLM for parallel verification to ensure statistical consistency. However, periodic parallel calls to the target LLM for verification prevent SD from achieving even lower

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).