← all papers · overview

Progressively Label Enhancement For Large Language Model Alignment

Abstract

Large Language Models (LLM) alignment aims to prevent models from producing content that misaligns with human expectations, which can lead to ethical and legal concerns. In the last few years, Reinforcement Learning from Human Feedback (RLHF) has been the most prominent method for achieving alignment. Due to challenges in stability and scalability with RLHF stages, which arise from the complex int

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).