tau-2-bench
Emerging17papers using it
2025first seen
'Tau-2 Bench' is a dataset used to evaluate the performance of tool-use agents by providing a structured set of tasks that assess interaction dynamics and the effectiveness of various training strategies.
Papers using tau-2-bench (17)
- Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience LearningTowards General Agentic Intelligence via Environment ScalingPolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM AgentsAgent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM AgentsProper Scoring Rules for Agentic Uncertainty QuantificationSkillsInjector: Dynamic Skill Context Construction for LLM AgentsCurateEvo: Data-Curation Evolving for Agentic Post-TrainingMemGym: a Long-Horizon Memory Environment for LLM AgentsMAVEN: Improving Generalization in Agentic Tool CallingRobust Tool Use via Fission-GRPO: Learning to Recover from Execution ErrorsTopoCurate:Modeling Interaction Topology for Tool-Use Agent TrainingEnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RLToward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM AgentsStep 3.5 Flash: Open Frontier-level Intelligence With 11B Active ParametersAutoForge: Automated Environment Synthesis for Agentic Reinforcement LearningToolorchestra: Elevating Intelligence Via Efficient Model And Tool OrchestrationOn Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset