← all papers · overview

Efficient Heterogeneous Large Language Model Decoding With Model-attention Disaggregation

Abstract

Transformer-based large language models (LLMs) exhibit impressive performance in generative tasks but also introduce significant challenges in real-world serving due to inefficient use of the expensive, computation-optimized accelerators. Although disaggregated serving architectures have been proposed to split different phases of LLM inference, the efficiency of decoding phase is still low. This i

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).