← all papers · overview

Optimized Multi-token Joint Decoding With Auxiliary Model For LLM Inference

Abstract

Large language models (LLMs) have achieved remarkable success across diverse tasks, yet their inference processes are hindered by substantial time and energy demands due to single-token generation at each decoding step. While previous methods such as speculative decoding mitigate these inefficiencies by producing multiple tokens per step, each token is still generated by its single-token distribut

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).