← all papers · overview

Parallel Decoding Via Hidden Transfer For Lossless Large Language Model Acceleration

Abstract

Large language models (LLMs) have recently shown remarkable performance across a wide range of tasks. However, the substantial number of parameters in LLMs contributes to significant latency during model inference. This is particularly evident when utilizing autoregressive decoding methods, which generate one token in a single forward process, thereby not fully capitalizing on the parallel computi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).