← all papers · overview

Emergent Misalignment Is Easy, Narrow Misalignment Is Hard

Abstract

Finetuning large language models on narrowly harmful datasets can cause them to become emergently misaligned, giving stereotypically `evil' responses across diverse unrelated settings. Concerningly, a pre-registered survey of experts failed to predict this result, highlighting our poor understanding of the inductive biases governing learning and generalisation in LLMs. We use emergent misalignment

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).