← all papers · overview

Poser: Unmasking Alignment Faking Llms By Manipulating Their Internals

Abstract

Like a criminal under investigation, Large Language Models (LLMs) might pretend to be aligned while evaluated and misbehave when they have a good opportunity. Can current interpretability methods catch these 'alignment fakers?' To answer this question, we introduce a benchmark that consists of 324 pairs of LLMs fine-tuned to select actions in role-play scenarios. One model in each pair is consiste

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).