← all papers · overview

Diffusionattacker: Diffusion-driven Prompt Manipulation For LLM Jailbreak

Abstract

Large Language Models (LLMs) are susceptible to generating harmful content when prompted with carefully crafted inputs, a vulnerability known as LLM jailbreaking. As LLMs become more powerful, studying jailbreak methods is critical to enhancing security and aligning models with human values. Traditionally, jailbreak techniques have relied on suffix addition or prompt templates, but these methods s

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).