← all papers · overview

Sparse Mezo: Less Parameters For Better Performance In Zeroth-order LLM Fine-tuning

Abstract

While fine-tuning large language models (LLMs) for specific tasks often yields impressive results, it comes at the cost of memory inefficiency due to back-propagation in gradient-based training. Memory-efficient Zeroth-order (MeZO) optimizers, recently proposed to address this issue, only require forward passes during training, making them more memory-friendly. However, compared with exact gradien

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).