← all papers · overview

Automixq: Self-adjusting Quantization For High Performance Memory-efficient Fine-tuning

Abstract

Fine-tuning large language models (LLMs) under resource constraints is a significant challenge in deep learning. Low-Rank Adaptation (LoRA), pruning, and quantization are all effective methods for improving resource efficiency. However, combining them directly often results in suboptimal performance, especially with uniform quantization across all model layers. This is due to the complex, uneven i

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).