← all papers · overview

Large Language Model Sentinel: LLM Agent For Adversarial Purification

Abstract

Over the past two years, the use of large language models (LLMs) has advanced rapidly. While these LLMs offer considerable convenience, they also raise security concerns, as LLMs are vulnerable to adversarial attacks by some well-designed textual perturbations. In this paper, we introduce a novel defense technique named Large LAnguage MOdel Sentinel (LLAMOS), which is designed to enhance the adver

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).