Terminal-Bench 2.0
Emerging14papers using it
2026first seen
Warning: The leaderboard above is unofficial. The official leaderboard is https://www.tbench.ai/leaderboard/terminal-bench/2.0, in which entires are audited for correct configuration, results show which agent harness is used, and verified trajectories are publicly viewable. Warning: The dataset is a read-only mirror. T
Papers using Terminal-Bench 2.0 (12)
- Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent SkillsHarnessBridge: Learnable Bidirectional Controller for LLM Agent HarnessRemember When It Matters: Proactive Memory Agent for Long-Horizon AgentsLiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language AgentsTerminal-World: Scaling Terminal-Agent Environments via Agent SkillsECHO: Terminal Agents Learn World Models for FreeCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentsWhat Makes Interaction Trajectories Effective for Training Terminal Agents?APEX: Adaptive Principle EXtraction A Three-Layer Self-Evolution Framework for Production AI AgentsSEAGym: An Evaluation Environment for Self-Evolving LLM AgentsTerminal-bench: Benchmarking Agents On Hard, Realistic Tasks In Command Line InterfacesSkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution