← all papers · overview

Super(ficial)-alignment: Strong Models May Deceive Weak Models In Weak-to-strong Generalization

Abstract

Superalignment, where humans act as weak supervisors for superhuman models, has become a crucial problem with the rapid development of Large Language Models (LLMs). Recent work has preliminarily studied this problem by using weak models to supervise strong models, and discovered that weakly supervised strong students can consistently outperform weak teachers towards the alignment target, leading t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).