← all papers · overview

Rethinking How To Evaluate Language Model Jailbreak

Abstract

Large language models (LLMs) have become increasingly integrated with various applications. To ensure that LLMs do not generate unsafe responses, they are aligned with safeguards that specify what content is restricted. However, such alignment can be bypassed to produce prohibited content using a technique commonly referred to as jailbreak. Different systems have been proposed to perform the jailb

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).