← all papers · overview

APB: Accelerating Distributed Long-context Inference By Passing Compressed Context Blocks Across Gpus

Abstract

While long-context inference is crucial for advancing large language model (LLM) applications, its prefill speed remains a significant bottleneck. Current approaches, including sequence parallelism strategies and compute reduction through approximate attention mechanisms, still fall short of delivering optimal inference efficiency. This hinders scaling the inputs to longer sequences and processing

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).