← all papers · overview

Exploiting LLM Quantization

Abstract

Quantization leverages lower-precision weights to reduce the memory usage of large language models (LLMs) and is a key technique for enabling their deployment on commodity hardware. While LLM quantization's impact on utility has been extensively explored, this work for the first time studies its adverse effects from a security perspective. We reveal that widely used quantization methods can be exp

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).