← all papers · overview

Contrastive Policy Gradient: Aligning Llms On Sequence-level Scores In A Supervised-friendly Fashion

Abstract

Reinforcement Learning (RL) has been used to finetune Large Language Models (LLMs) using a reward model trained from preference data, to better align with human judgment. The recently introduced direct alignment methods, which are often simpler, more stable, and computationally lighter, can more directly achieve this. However, these approaches cannot optimize arbitrary rewards, and the preference-

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).