← all papers · overview

Provable Last-iterate Convergence For Multi-objective Safe LLM Alignment Via Optimistic Primal-dual

Abstract

Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. While RLHF with expected reward constraints can be formulated as a primal-dual optimization problem, standard primal-dual methods only guarantee convergence with a distributional policy where the saddle-point problem is in convex-concave form. Moreover, standa

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).