← all papers · overview

Mrj-agent: An Effective Jailbreak Agent For Multi-round Dialogue

Abstract

Large Language Models (LLMs) demonstrate outstanding performance in their reservoir of knowledge and understanding capabilities, but they have also been shown to be prone to illegal or unethical reactions when subjected to jailbreak attacks. To ensure their responsible deployment in critical applications, it is crucial to understand the safety capabilities and vulnerabilities of LLMs. Previous wor

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).