← all papers · overview

Steering Frozen Llms: Adaptive Social Alignment Via Online Prompt Routing

Abstract

Large language models (LLMs) are typically governed by post-training alignment (e.g., RLHF or DPO), which yields a largely static policy during deployment and inference. However, real-world safety is a full-lifecycle problem: static defenses degrade against evolving jailbreak behaviors, and fixed weights cannot adapt to pluralistic, time-varying safety norms. This motivates inference-time governan

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).