← all papers · overview

Bifurcated Attention: Accelerating Massively Parallel Decoding With Shared Prefixes In Llms

Abstract

This study introduces bifurcated attention, a method designed to enhance language model inference in shared-context batch decoding scenarios. Our approach addresses the challenge of redundant memory IO costs, a critical factor contributing to latency in high batch sizes and extended context lengths. Bifurcated attention achieves this by strategically dividing the attention mechanism during increme

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).