← all papers · overview

Acttail: Global Activation Sparsity In Large Language Models

Abstract

Activation sparsity is a promising approach for accelerating large language model (LLM) inference by reducing computation and memory movement. However, existing activation sparsity methods typically apply uniform sparsity across projections, ignoring the heterogeneous statistical properties of Transformer weights and thereby amplifying performance degradation. In this paper, we propose ActTail, a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).