← all papers · overview

I-LLM: Efficient Integer-only Inference For Fully-quantized Low-bit Large Language Models

Abstract

Post-training quantization (PTQ) serves as a potent technique to accelerate the inference of large language models (LLMs). Nonetheless, existing works still necessitate a considerable number of floating-point (FP) operations during inference, including additional quantization and de-quantization, as well as non-linear operators such as RMSNorm and Softmax. This limitation hinders the deployment of

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).