← all papers · overview

Prompt Attack Detection With Llm-as-a-judge And Mixture-of-models

Abstract

Prompt attacks, including jailbreaks and prompt injections, pose a critical security risk to Large Language Model (LLM) systems. In production, guardrails must mitigate these attacks under strict low-latency constraints, resulting in a deployment gap in which lightweight classifiers and rule-based systems struggle to generalize under distribution shift, while high-capacity LLM-based judges remain

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).