← all papers · overview

Pipeinfer: Accelerating LLM Inference Using Asynchronous Pipelined Speculation

Abstract

Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from CPU speculative execution. These techniques reduce bottlenecks associated with memory bandwidth, but also increase end-to-end latency per inference run, requiring high speculation acceptance rates to improve performance.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).