← all papers · overview

Ruquant: Towards Refining Uniform Quantization For Large Language Models

Abstract

The increasing size and complexity of large language models (LLMs) have raised significant challenges in deployment efficiency, particularly under resource constraints. Post-training quantization (PTQ) has emerged as a practical solution by compressing models without requiring retraining. While existing methods focus on uniform quantization schemes for both weights and activations, they often suff

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).