← all papers · overview

Layer-level Self-exposure And Patch: Affirmative Token Mitigation For Jailbreak Attack Defense

Abstract

As large language models (LLMs) are increasingly deployed in diverse applications, including chatbot assistants and code generation, aligning their behavior with safety and ethical standards has become paramount. However, jailbreak attacks, which exploit vulnerabilities to elicit unintended or harmful outputs, threaten LLMs' safety significantly. In this paper, we introduce Layer-AdvPatcher, a nov

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).