← all papers · overview

Hardware-aware Parallel Prompt Decoding For Memory-efficient Acceleration Of LLM Inference

Abstract

The auto-regressive decoding of Large Language Models (LLMs) results in significant overheads in their hardware performance. While recent research has investigated various speculative decoding techniques for multi-token generation, these efforts have primarily focused on improving processing speed such as throughput. Crucially, they often neglect other metrics essential for real-life deployments,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).