← all papers · overview

Censored Llms As A Natural Testbed For Secret Knowledge Elicitation

Abstract

Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answers truthfully -- and lie detection -- classifying whether a given response is false. Prior work evaluates such methods on models specifically trained to lie or conceal information, but these artificial constructions may

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).