← all papers · overview

PLD+: Accelerating LLM Inference By Leveraging Language Model Artifacts

Abstract

To reduce the latency associated with autoretrogressive LLM inference, speculative decoding has emerged as a novel decoding paradigm, where future tokens are drafted and verified in parallel. However, the practical deployment of speculative decoding is hindered by its requirements for additional computational resources and fine-tuning, which limits its out-of-the-box usability. To address these ch

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).