← all papers · overview

Accelerating Greedy Coordinate Gradient And General Prompt Optimization Via Probe Sampling

Abstract

Safety of Large Language Models (LLMs) has become a critical issue given their rapid progresses. Greedy Coordinate Gradient (GCG) is shown to be effective in constructing adversarial prompts to break the aligned LLMs, but optimization of GCG is time-consuming. To reduce the time cost of GCG and enable more comprehensive studies of LLM safety, in this work, we study a new algorithm called \(\texttt

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).