← all papers · overview

Iteralign: Iterative Constitutional Alignment Of Large Language Models

Abstract

With the rapid development of large language models (LLMs), aligning LLMs with human values and societal norms to ensure their reliability and safety has become crucial. Reinforcement learning with human feedback (RLHF) and Constitutional AI (CAI) have been proposed for LLM alignment. However, these methods require either heavy human annotations or explicitly pre-defined constitutions, which are l

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).