← all papers · overview

Online Merging Optimizers For Boosting Rewards And Mitigating Tax In Alignment

Abstract

Effectively aligning Large Language Models (LLMs) with human-centric values while preventing the degradation of abilities acquired through Pre-training and Supervised Fine-tuning (SFT) poses a central challenge in Reinforcement Learning from Human Feedback (RLHF). In this paper, we first discover that interpolating RLHF and SFT model parameters can adjust the trade-off between human preference and

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).