← all papers · overview

Amphista: Bi-directional Multi-head Decoding For Accelerating LLM Inference

Abstract

Large Language Models (LLMs) inherently use autoregressive decoding, which lacks parallelism in inference and results in significantly slow inference speed. While methods such as Medusa constructs parallelized heads, they lack adequate information interaction across different prediction positions. To overcome this limitation, we introduce Amphista, an enhanced speculative decoding framework that b

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).