← all papers · overview

Linear Predictability Of Attention Heads In Large Language Models

Abstract

Large language model (LLM) inference is increasingly bottlenecked by the Key-Value (KV) cache, yet the fine-grained structure of attention-head activations remains poorly understood. We show that pretrained Transformers exhibit a pervasive inter-head linear structure: for a given token, the Query, Key, and Value (QKV) vectors of an attention head can often be reconstructed as a linear combination

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).