← all papers · overview

Optimizing Safe And Aligned Language Generation: A Multi-objective GRPO Approach

Abstract

Aligning large language models (LLMs) with human values and safety constraints is challenging, especially when objectives like helpfulness, truthfulness, and avoidance of harm conflict. Reinforcement Learning from Human Feedback (RLHF) has achieved notable success in steering models, but is complex and can be unstable. Recent approaches such as Direct Preference Optimization (DPO) simplify prefere

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).