← all papers · overview

Magnetic Preference Optimization: Achieving Last-iterate Convergence For Language Model Alignment

Abstract

Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF), self-play not only boosts Large Language Model (LLM) performance but also overcomes the limitations of traditional Bradley-Terry (BT) model assumptions by finding the Nash equilibrium (NE) of a preference-based, two-play

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).