AgentDojo
Emerging8papers using it
2025first seen
AgentDojo is a benchmark that evaluates the effectiveness of monitoring protocols against indirect prompt injection attacks on AI agents.
Papers using AgentDojo (8)
- AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM AgentsAgentdyn: A Dynamic Open-ended Benchmark For Evaluating Prompt Injection Attacks Of Real-world Agent Security SystemAgent-Sentry: Bounding LLM Agents via Execution ProvenanceLearning to Inject: Automated Prompt Injection via Reinforcement LearningMitigating Indirect Prompt Injection via Instruction-Following Intent AnalysisIndirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?CommandSans: Securing AI Agents with Surgical Precision Prompt SanitizationPromptArmor: Simple yet Effective Prompt Injection Defenses