← all papers · overview

Exploring The Impact Of Low-rank Adaptation On The Performance, Efficiency, And Regularization Of RLHF

Abstract

During the last stage of RLHF, a large language model is aligned to human intents via PPO training, a process that generally requires large-scale computational resources. In this technical report, we empirically investigate an efficient implementation of RLHF using low-rank adaptation (LoRA), which allows us to align the LLaMA 7B checkpoint on the Alpaca dataset using only two A100 GPUs instead of

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).