Terminal-Bench
Emerging15papers using it
2025first seen
Terminal-Bench Dataset This dataset contains tasks from Terminal-Bench, a benchmark for evaluating AI agents in real terminal environments. Each task is packaged as a complete, self-contained archive that preserves the exact directory structure, binary files, Docker configurations, and test scripts needed for faithful
Papers using Terminal-Bench (9)
- Qwen3-Coder-Next Technical ReportCLI-Gym: Scalable CLI Task Generation via Agentic Environment InversionADK Arena: Evaluating Agent Development Kits via LLM-as-a-DeveloperCRANE: Constrained Reasoning Injection for Code Agents via Nullspace EditingR2V Agent: Teaching SLMs When to Ask for HelpFrom Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to PythonToward Scalable Terminal Task Synthesis via Skill GraphsAOrchestra: Automating Sub-Agent Creation for Agentic OrchestrationSkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent