← all papers · overview

Efficient And Adaptable Detection Of Malicious LLM Prompts Via Bootstrap Aggregation

Abstract

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding, reasoning, and generation. However, these systems remain susceptible to malicious prompts that induce unsafe or policy-violating behavior through harmful requests, jailbreak techniques, and prompt injection attacks. Existing defenses face fundamental limitations: black-box moderation APIs offe

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).