← all papers · overview

Llmscan: Causal Scan For LLM Misbehavior Detection

Abstract

Despite the success of Large Language Models (LLMs) across various fields, their potential to generate untruthful, biased and harmful responses poses significant risks, particularly in critical applications. This highlights the urgent need for systematic methods to detect and prevent such misbehavior. While existing approaches target specific issues such as harmful responses, this work introduces

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).