← all papers · overview

Fira: Can We Achieve Full-rank Training Of Llms Under Low-rank Constraint?

Abstract

Low-rank training has emerged as a promising approach for reducing memory usage in training Large Language Models (LLMs). Previous methods either rely on decomposing weight matrices (e.g., LoRA), or seek to decompose gradient matrices (e.g., GaLore) to ensure reduced memory consumption. However, both of them constrain the training in a low-rank subspace, thus inevitably leading to sub-optimal perf

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).