← all papers · overview

DHA: Learning Decoupled-head Attention From Transformer Checkpoints Via Adaptive Heads Fusion

Abstract

Large language models (LLMs) with billions of parameters demonstrate impressive performance. However, the widely used Multi-Head Attention (MHA) in LLMs incurs substantial computational and memory costs during inference. While some efforts have optimized attention mechanisms by pruning heads or sharing parameters among heads, these methods often lead to performance degradation or necessitate subst

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).