← all papers · overview

Do Prompts Guarantee Safety? Mitigating Toxicity From LLM Generations Through Subspace Intervention

Abstract

Large Language Models (LLMs) are powerful text generators, yet they can produce toxic or harmful content even when given seemingly harmless prompts. This presents a serious safety challenge and can cause real-world harm. Toxicity is often subtle and context-dependent, making it difficult to detect at the token level or through coarse sentence-level signals. Moreover, efforts to mitigate toxicity o

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).