← all papers · overview

Sandwich Attack: Multi-language Mixture Adaptive Attack On Llms

Abstract

Large Language Models (LLMs) are increasingly being developed and applied, but their widespread use faces challenges. These include aligning LLMs' responses with human values to prevent harmful outputs, which is addressed through safety training methods. Even so, bad actors and malicious users have succeeded in attempts to manipulate the LLMs to generate misaligned responses for harmful questions

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).