← all papers · overview

Eraser: Jailbreaking Defense In Large Language Models Via Unlearning Harmful Knowledge

Abstract

Jailbreaking attacks can enable Large Language Models (LLMs) to bypass the safeguard and generate harmful content. Existing jailbreaking defense methods have failed to address the fundamental issue that harmful knowledge resides within the model, leading to potential jailbreak risks for LLMs. In this paper, we propose a novel defense method called Eraser, which mainly includes three goals: unlearn

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).