← all papers · overview

Attention Head Entropy Of Llms Predicts Answer Correctness

Abstract

Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-judge approaches risk introducing hidden errors. Recent white-box methods detect contextual hallucinations using model internals, focusing on the localization of the attention mass, but two questions remain open: do these

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).