← all papers · overview

Truth-value Judgment In Language Models: 'truth Directions' Are Context Sensitive

Abstract

Recent work has demonstrated that the latent spaces of large language models (LLMs) contain directions predictive of the truth of sentences. Multiple methods recover such directions and build probes that are described as uncovering a model's "knowledge" or "beliefs". We investigate this phenomenon, looking closely at the impact of context on the probes. Our experiments establish where in the LLM t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).