← all papers · overview

Aligning Large Language Models From Self-reference AI Feedback With One General Principle

Abstract

In aligning large language models (LLMs), utilizing feedback from existing advanced AI rather than humans is an important method to scale supervisory signals. However, it is highly challenging for AI to understand human intentions and societal values, and provide accurate preference feedback based on these. Current AI feedback methods rely on powerful LLMs, carefully designed specific principles t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).