← all papers · overview

When Models Examine Themselves: Vocabulary-activation Correspondence In Self-referential Processing

Abstract

Large language models produce rich introspective language when prompted for self-examination, but whether this language reflects internal computation or sophisticated confabulation has remained unclear. We show that self-referential vocabulary tracks concurrent activation dynamics, and that this correspondence is specific to self-referential processing. We introduce the Pull Methodology, a protoco

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).