← all papers · overview

Jigsaw Puzzles: Splitting Harmful Questions To Jailbreak Large Language Models

Abstract

Large language models (LLMs) have exhibited outstanding performance in engaging with humans and addressing complex questions by leveraging their vast implicit knowledge and robust reasoning capabilities. However, such models are vulnerable to jailbreak attacks, leading to the generation of harmful responses. Despite recent research on single-turn jailbreak strategies to facilitate the development

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).