← all papers · overview

Effective And Efficient Jailbreaks Of Black-box Llms With Cross-behavior Attacks

Abstract

Despite recent advancements in Large Language Models (LLMs) and their alignment, they can still be jailbroken, i.e., harmful and toxic content can be elicited from them. While existing red-teaming methods have shown promise in uncovering such vulnerabilities, these methods struggle with limited success and high computational and monetary costs. To address this, we propose a black-box Jailbreak met

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).