← all papers · overview

Specdec++: Boosting Speculative Decoding Via Adaptive Candidate Lengths

Abstract

Speculative decoding reduces the inference latency of a target large language model via utilizing a smaller and faster draft model. Its performance depends on a hyperparameter K -- the candidate length, i.e., the number of candidate tokens for the target model to verify in each round. However, previous methods often use simple heuristics to choose K, which may result in sub-optimal performance. We

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).