Claude Code
Emerging6papers using it
2026first seen
The 'Claude Code' dataset/benchmark is used to evaluate the performance of LLM-driven automated penetration testing agents against intelligent defense mechanisms in a realistic, turn-based attacker-defender-judge framework.