← all papers · overview

Gradient-controlled Decoding: A Safety Guardrail For Llms With Dual-anchor Steering

Abstract

Large language models (LLMs) remain susceptible to jailbreak and direct prompt-injection attacks, yet the strongest defensive filters frequently over-refuse benign queries and degrade user experience. Previous work on jailbreak and prompt injection detection such as GradSafe, detects unsafe prompts with a single "accept all" anchor token, but its threshold is brittle and it offers no deterministic

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).