← all papers · overview

All In How You Ask For It: Simple Black-box Method For Jailbreak Attacks

Abstract

Large Language Models (LLMs), such as ChatGPT, encounter `jailbreak' challenges, wherein safeguards are circumvented to generate ethically harmful prompts. This study introduces a straightforward black-box method for efficiently crafting jailbreak prompts, addressing the significant complexity and computational costs associated with conventional methods. Our technique iteratively transforms harmfu

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).