Despite their impressive capabilities, large language models (LLMs) frequently generate hallucinations. Previous work shows that their internal states encode rich signals of truthfulness, yet the origins and mechanisms of these signals remain unclear. In this paper, we demonstrate that truthfulness cues arise from two distinct information pathwa
Related papers
Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).