← all papers · overview

Curveball Steering: The Right Direction To Steer Isn't Always Linear

Abstract

Activation steering is a widely used approach for controlling large language model (LLM) behavior by intervening on internal representations. Existing methods largely rely on the Linear Representation Hypothesis, assuming behavioral attributes can be manipulated using global linear directions. In practice, however, such linear interventions often behave inconsistently. We question this assumption

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).