← all papers · overview

Self-distillation For Multi-token Prediction

Abstract

As Large Language Models (LLMs) scale up, inference efficiency becomes a critical bottleneck. Multi-Token Prediction (MTP) could accelerate LLM inference by predicting multiple future tokens in parallel. However, existing MTP approaches still face two challenges: limited acceptance rates of MTP heads, and difficulties in jointly training multiple MTP heads. Therefore, we propose MTP-D, a simple ye

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).