← all papers · overview

Decoding Human Preferences In Alignment: An Improved Approach To Inverse Constitutional AI

Abstract

Traditional methods for aligning Large Language Models (LLMs), such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on implicit principles, limiting interpretability. Constitutional AI (CAI) offers an explicit, rule-based framework for guiding LLM alignment. Building on this, we refine the Inverse Constitutional AI (ICAI) algorithm, which extract

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).