← all papers · overview

Odesteer: A Unified Ode-based Steering Framework For LLM Alignment

Abstract

Activation steering, or representation engineering, offers a lightweight approach to align large language models (LLMs) by manipulating their internal activations at inference time. However, current methods suffer from two key limitations: (i) the lack of a unified theoretical framework for guiding the design of steering directions, and (ii) an over-reliance on one-step steering that fail to captu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).