← all papers · overview

Safeguarding Large Language Models In Real-time With Tunable Safety-performance Trade-offs

Abstract

Large Language Models (LLMs) have been shown to be susceptible to jailbreak attacks, or adversarial attacks used to illicit high risk behavior from a model. Jailbreaks have been exploited by cybercriminals and blackhat actors to cause significant harm, highlighting the critical need to safeguard widely-deployed models. Safeguarding approaches, which include fine-tuning models or having LLMs "self-

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).