tau-2-bench
Emerging10papers using it
2025first seen
The 'tau-2-bench' is a benchmark that evaluates the performance of models in orchestrating multi-step tool calls within realistic stateful execution environments.
Papers using tau-2-bench (10)
- Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live EnvironmentsEnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RLTopoCurate:Modeling Interaction Topology for Tool-Use Agent TrainingKAT-Coder-V2 Technical ReportStep 3.5 Flash: Open Frontier-Level Intelligence with 11B Active ParametersRobust Tool Use via Fission-GRPO: Learning to Recover from Execution ErrorsFrom Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using AgentsAutoForge: Automated Environment Synthesis for Agentic Reinforcement LearningToolOrchestra: Elevating Intelligence via Efficient Model and Tool OrchestrationKimi K2: Open Agentic Intelligence