← all papers · overview

Probing The Limits Of The Lie Detector Approach To LLM Deception

Abstract

Mechanistic approaches to deception in large language models (LLMs) often rely on "lie detectors", that is, truth probes trained to identify internal representations of model outputs as false. The lie detector approach to LLM deception implicitly assumes that deception is coextensive with lying. This paper challenges that assumption. It experimentally investigates whether LLMs can deceive without

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).