← all papers · overview

RED QUEEN: Safeguarding Large Language Models Against Concealed Multi-turn Jailbreaking

Abstract

The rapid progress of Large Language Models (LLMs) has opened up new opportunities across various domains and applications; yet it also presents challenges related to potential misuse. To mitigate such risks, red teaming has been employed as a proactive security measure to probe language models for harmful outputs via jailbreak attacks. However, current jailbreak attack approaches are single-turn

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).