← all papers · overview

SDSAT: Accelerating LLM Inference Through Speculative Decoding With Semantic Adaptive Tokens

Abstract

We propose an acceleration scheme for large language models (LLMs) through Speculative Decoding with Semantic Adaptive Tokens (SDSAT). The primary objective of this design is to enhance the LLM model's ability to generate draft tokens more accurately without compromising the model's accuracy. The core strategies involve: 1) Fine-tune the model by incorporating semantic adaptive tokens that possess

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).