← all papers · overview

Hylra: Hybrid Layer Reuse Attention For Efficient Long-context Inference

Abstract

Long-context inference in Large Language Models (LLMs) is bottlenecked by the quadratic computation complexity of attention and the substantial memory footprint of Key-Value (KV) caches. While existing sparse attention mechanisms attempt to mitigate this by exploiting inherent sparsity, they often rely on rigid patterns or aggressive pruning, failing to achieve an optimal balance between efficienc

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).