← all papers · overview

PRISM: Parametrically Refactoring Inference For Speculative Sampling Draft Models

Abstract

Large Language Models (LLMs), constrained by their auto-regressive nature, suffer from slow decoding. Speculative decoding methods have emerged as a promising solution to accelerate LLM decoding, attracting attention from both systems and AI research communities. Recently, the pursuit of better draft quality has driven a trend toward parametrically larger draft models, which inevitably introduces

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).