← all papers · overview

Harnessing Task Overload For Scalable Jailbreak Attacks On Large Language Models

Abstract

Large Language Models (LLMs) remain vulnerable to jailbreak attacks that bypass their safety mechanisms. Existing attack methods are fixed or specifically tailored for certain models and cannot flexibly adjust attack strength, which is critical for generalization when attacking models of various sizes. We introduce a novel scalable jailbreak attack that preempts the activation of an LLM's safety p

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).