← all papers · overview

L4Q: Parameter Efficient Quantization-aware Fine-tuning On Large Language Models

Abstract

Due to the high memory and computational costs associated with large language models (LLMs), model compression techniques such as quantization, which reduces inference costs, and parameter-efficient fine-tuning (PEFT) methods like Low-Rank Adaptation (LoRA), which reduce training costs, have gained significant popularity. This trend has spurred active research into quantization-aware PEFT techniqu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).