← all papers · overview

SDFP: Speculative Decoding With Fit-pruned Models For Training-free And Plug-and-play LLM Acceleration

Abstract

Large language models (LLMs) underpin interactive multimedia applications such as captioning, retrieval, recommendation, and creative content generation, yet their autoregressive decoding incurs substantial latency. Speculative decoding reduces latency using a lightweight draft model, but deployment is often limited by the cost and complexity of acquiring, tuning, and maintaining an effective draf

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).