← all papers · overview

Faster-gcg: Efficient Discrete Optimization Jailbreak Attacks Against Aligned Large Language Models

Abstract

Aligned Large Language Models (LLMs) have demonstrated remarkable performance across various tasks. However, LLMs remain susceptible to jailbreak adversarial attacks, where adversaries manipulate prompts to elicit malicious responses that aligned LLMs should have avoided. Identifying these vulnerabilities is crucial for understanding the inherent weaknesses of LLMs and preventing their potential m

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).