← all papers · overview

Beyond Refusal: Probing The Limits Of Agentic Self-correction For Semantic Sensitive Information

Abstract

While defenses for structured PII are mature, Large Language Models (LLMs) pose a new threat: Semantic Sensitive Information (SemSI), where models infer sensitive identity attributes, generate reputation-harmful content, or hallucinate potentially wrong information. The capacity of LLMs to self-regulate these complex, context-dependent sensitive information leaks without destroying utility remains

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).