← all papers · overview

PARDEN, Can You Repeat That? Defending Against Jailbreaks Via Repetition

Abstract

Large language models (LLMs) have shown success in many natural language processing tasks. Despite rigorous safety alignment processes, supposedly safety-aligned LLMs like Llama 2 and Claude 2 are still susceptible to jailbreaks, leading to security risks and abuse of the models. One option to mitigate such risks is to augment the LLM with a dedicated "safeguard", which checks the LLM's inputs or

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).