Terminal-Bench
Emerging6papers using it
2025first seen
Terminal-Bench Dataset This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful
Papers using Terminal-Bench (6)
- ADK Arena: Evaluating Agent Development Kits via LLM-as-a-DeveloperCVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security VulnerabilityFrom Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to PythonQwen3-Coder-Next Technical ReportComposer 2 Technical ReportSkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent