← all papers · overview

BOSCH: Black-box Binary Optimization For Short-context Attention-head Selection In Llms

Abstract

Post-training hybridization of large language models (LLMs) often replaces quadratic self-attention with sliding-window attention (SWA) to reduce KV cache usage and improve latency. Existing hybridization schemes are typically defined either at the layer level (e.g., interleaving) or at the head level via static rankings from local to global. Layer-level schemes ignore that local and global depend

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).