← all papers · overview

DETAM: Defending Llms Against Jailbreak Attacks Via Targeted Attention Modification

Abstract

With the widespread adoption of Large Language Models (LLMs), jailbreak attacks have become an increasingly pressing safety concern. While safety-aligned LLMs can effectively defend against normal harmful queries, they remain vulnerable to such attacks. Existing defense methods primarily rely on fine-tuning or input modification, which often suffer from limited generalization and reduced utility.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).