← all papers · overview

Efficiently Aligning Draft Models Via Parameter- And Data-efficient Adaptation

Abstract

Speculative decoding accelerates LLM inference but suffers from performance degradation when target models are fine-tuned for specific domains. A naive solution is to retrain draft models for every target model, which is costly and inefficient. To address this, we introduce a parameter- and data-efficient framework named Efficient Draft Adaptation, abbreviated as EDA, for efficiently adapting draf

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).