← all papers · overview

Stochastic Monkeys At Play: Random Augmentations Cheaply Break LLM Safety Alignment

Abstract

Safety alignment of Large Language Models (LLMs) has recently become a critical objective of model developers. In response, a growing body of work has been investigating how safety alignment can be bypassed through various jailbreaking methods, such as adversarial attacks. However, these jailbreak methods can be rather costly or involve a non-trivial amount of creativity and effort, introducing th

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).