← all papers · overview

Language Models Fail To Introspect About Their Knowledge Of Language

Abstract

There has been recent interest in whether large language models (LLMs) can introspect about their own internal states. Such abilities would make LLMs more interpretable, and also validate the use of standard introspective methods in linguistics to evaluate grammatical knowledge in models (e.g., asking "Is this sentence grammatical?"). We systematically investigate emergent introspection across 21

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).