← all papers · overview

Advancing Adversarial Suffix Transfer Learning On Aligned Large Language Models

Abstract

Language Language Models (LLMs) face safety concerns due to potential misuse by malicious users. Recent red-teaming efforts have identified adversarial suffixes capable of jailbreaking LLMs using the gradient-based search algorithm Greedy Coordinate Gradient (GCG). However, GCG struggles with computational inefficiency, limiting further investigations regarding suffix transferability and scalabili

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).