← all papers · overview

Evolving Jailbreaks: Automated Multi-objective Long-tail Attacks On Large Language Models

Abstract

Large Language Models (LLMs) have been widely deployed, especially through free Web-based applications that expose them to diverse user-generated inputs, including those from long-tail distributions such as low-resource languages and encrypted private data. This open-ended exposure increases the risk of jailbreak attacks that undermine model safety alignment. While recent studies have shown that l

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).