← all papers · overview

Enhancing Computation Efficiency In Large Language Models Through Weight And Activation Quantization

Abstract

Large Language Models (LLMs) are proficient in natural language processing tasks, but their deployment is often restricted by extensive parameter sizes and computational demands. This paper focuses on post-training quantization (PTQ) in LLMs, specifically 4-bit weight and 8-bit activation (W4A8) quantization, to enhance computational efficiency -- a topic less explored compared to weight-only quan

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).