← all papers · overview

Evopress: Accurate Dynamic Model Compression Via Evolutionary Search

Abstract

The high computational costs of large language models (LLMs) have led to a flurry of research on LLM compression, via methods such as quantization, sparsification, or structured pruning. A new frontier in this area is given by dynamic, non-uniform compression methods, which adjust the compression levels (e.g., sparsity) per-block or even per-layer in order to minimize accuracy loss, while guarante

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).