← all papers · overview

Thanos: A Block-wise Pruning Algorithm For Efficient Large Language Model Compression

Abstract

This paper presents Thanos, a novel weight-pruning algorithm designed to reduce the memory footprint and enhance the computational efficiency of large language models (LLMs) by removing redundant weights while maintaining accuracy. Thanos introduces a block-wise pruning strategy with adaptive masks that dynamically adjust to weight importance, enabling flexible sparsity patterns and structured for

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).