← all papers · overview

Weakly Supervised Detection Of Hallucinations In LLM Activations

Abstract

We propose an auditing method to identify whether a large language model (LLM) encodes patterns such as hallucinations in its internal states, which may propagate to downstream tasks. We introduce a weakly supervised auditing technique using a subset scanning approach to detect anomalous patterns in LLM activations from pre-trained models. Importantly, our method does not need knowledge of the typ

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).