← all papers · overview

Pv-tuning: Beyond Straight-through Estimation For Extreme LLM Compression

Abstract

There has been significant interest in "extreme" compression of large language models (LLMs), i.e., to 1-2 bits per parameter, which allows such models to be executed efficiently on resource-constrained devices. Existing work focused on improved one-shot quantization techniques and weight representations; yet, purely post-training approaches are reaching diminishing returns in terms of the accurac

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).