← all papers · overview

Interpretable LLM Guardrails Via Sparse Representation Steering

Abstract

Large language models (LLMs) exhibit impressive capabilities in generation tasks but are prone to producing harmful, misleading, or biased content, posing significant ethical and safety concerns. To mitigate such risks, representation engineering, which steer model behavior toward desired attributes by injecting carefully designed steering vectors into LLM's representations at inference time, has

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).