← all papers · overview

Aligning Large Language Models By On-policy Self-judgment

Abstract

Existing approaches for aligning large language models with human preferences face a trade-off that requires a separate reward model (RM) for on-policy learning. In this paper, we present a novel alignment framework, SELF-JUDGE that (1) does on-policy learning and 2) is parameter efficient, as it does not require an additional RM for evaluating the samples for on-policy learning. To this end, we p

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).