← all papers · overview

Training Domain Draft Models For Speculative Decoding: Best Practices And Insights

Abstract

Speculative decoding is an effective method for accelerating inference of large language models (LLMs) by employing a small draft model to predict the output of a target model. However, when adapting speculative decoding to domain-specific target models, the acceptance rate of the generic draft model drops significantly due to domain shift. In this work, we systematically investigate knowledge dis

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).