← all papers · overview

Adazeta: Adaptive Zeroth-order Tensor-train Adaption For Memory-efficient Large Language Models Fine-tuning

Abstract

Fine-tuning large language models (LLMs) has achieved remarkable performance across various natural language processing tasks, yet it demands more and more memory as model sizes keep growing. To address this issue, the recently proposed Memory-efficient Zeroth-order (MeZO) methods attempt to fine-tune LLMs using only forward passes, thereby avoiding the need for a backpropagation graph. However, s

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).