← all papers · overview

Bypassing Prompt Injection Detectors Through Evasive Injections

Abstract

Large language models (LLMs) are increasingly used in interactive and retrieval-augmented systems, but they remain vulnerable to prompt injection attacks, where injected secondary prompts force the model to deviate from the user's instructions to execute a potentially malicious task defined by the adversary. Recent work shows that ML models trained on activation shifts from LLMs' hidden layers can

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).