← all datasets

Terminal-Bench 2

Emerging
4papers using it
2026first seen

'Terminal-Bench 2' is a dataset used to evaluate the performance and capabilities of AI agents and large language models by providing a collection of benchmark tasks that can be systematically audited for issues such as ambiguous task design and execution environment conflicts.

Papers using Terminal-Bench 2 (3)

Terminal-Bench 2 dataset β€” papers, benchmarks & downloads Β· AI Agents