← all papers · overview

Second-order Fine-tuning Without Pain For Llms:a Hessian Informed Zeroth-order Optimizer

Abstract

Fine-tuning large language models (LLMs) with classic first-order optimizers entails prohibitive GPU memory due to the backpropagation process. Recent works have turned to zeroth-order optimizers for fine-tuning, which save substantial memory by using two forward passes. However, these optimizers are plagued by the heterogeneity of parameter curvatures across different dimensions. In this work, we

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).