← all papers · overview

Theory Of Mind And Self-attributions Of Mentality Are Dissociable In Llms

Abstract

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to experience emotions. We investigate whether suppressing mind-attribution tendencies degrades intimately related socio-cognitive abilities such as Theory of Mind (ToM). Through safety ablation and mechanistic analyses of r

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).