← all papers · overview

Direct Alignment Of Draft Model For Speculative Decoding With Chat-fine-tuned Llms

Abstract

Text generation with Large Language Models (LLMs) is known to be memory bound due to the combination of their auto-regressive nature, huge parameter counts, and limited memory bandwidths, often resulting in low token rates. Speculative decoding has been proposed as a solution for LLM inference acceleration. However, since draft models are often unavailable in the modern open-source LLM families, e

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).