← all papers · overview

Cognitive Overload: Jailbreaking Large Language Models With Overloaded Logical Thinking

Abstract

While large language models (LLMs) have demonstrated increasing power, they have also given rise to a wide range of harmful behaviors. As representatives, jailbreak attacks can provoke harmful or unethical responses from LLMs, even after safety alignment. In this paper, we investigate a novel category of jailbreak attacks specifically designed to target the cognitive structure and processes of LLM

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).