← all papers · overview

Provable Defense Framework For LLM Jailbreaks Via Noise-augumented Alignment

Abstract

Large Language Models (LLMs) remain vulnerable to adaptive jailbreaks that easily bypass empirical defenses like GCG. We propose a framework for certifiable robustness that shifts safety guarantees from single-pass inference to the statistical stability of an ensemble. We introduce Certified Semantic Smoothing (CSS) via Stratified Randomized Ablation, a technique that partitions inputs into immuta

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).