← all papers · overview

Dynamic-width Speculative Beam Decoding For Efficient LLM Inference

Abstract

Large language models (LLMs) have shown outstanding performance across numerous real-world tasks. However, the autoregressive nature of these models makes the inference process slow and costly. Speculative decoding has emerged as a promising solution, leveraging a smaller auxiliary model to draft future tokens, which are then validated simultaneously by the larger model, achieving a speed-up of 1-

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).