← all papers · overview

Self-control Of LLM Behaviors By Compressing Suffix Gradient Into Prefix Controller

Abstract

We propose SelfControl, an inference-time model control method utilizing gradients to control the behavior of large language models (LLMs) without explicit human annotations. Given a desired behavior expressed in a natural language suffix string concatenated to the input prompt, SelfControl computes gradients of the LLM's self-evaluation of the suffix with respect to its latent representations. Th

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).