← all papers · overview

CATS: Contextually-aware Thresholding For Sparsity In Large Language Models

Abstract

Large Language Models (LLMs) have dramatically advanced AI applications, yet their deployment remains challenging due to their immense inference costs. Recent studies ameliorate the computational costs of LLMs by increasing their activation sparsity but suffer from significant performance degradation on downstream tasks. In this work, we introduce a new framework for sparsifying the activations of

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).