← all papers · overview

Steering The Censorship: Uncovering Representation Vectors For LLM "thought" Control

Abstract

Large language models (LLMs) have transformed the way we access information. These models are often tuned to refuse to comply with requests that are considered harmful and to produce responses that better align with the preferences of those who control the models. To understand how this "censorship" works. We use representation engineering techniques to study open-weights safety-tuned models. We p

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).