← all papers · overview

PACER: Blockwise Pre-verification For Speculative Decoding With Adaptive Length

Abstract

Speculative decoding (SD) is a powerful technique for accelerating the inference process of large language models (LLMs) without sacrificing accuracy. Typically, SD employs a small draft model to generate a fixed number of draft tokens, which are then verified in parallel by the target model. However, our experiments reveal that the optimal draft length varies significantly across different decodi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).