← all papers · overview

Seqar: Jailbreak Llms With Sequential Auto-generated Characters

Abstract

The widespread applications of large language models (LLMs) have brought about concerns regarding their potential misuse. Although aligned with human preference data before release, LLMs remain vulnerable to various malicious attacks. In this paper, we adopt a red-teaming strategy to enhance LLM safety and introduce SeqAR, a simple yet effective framework to design jailbreak prompts automatically.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).