ToolBench
Canonical12papers using it
2024first seen
ToolBench is a dataset used to evaluate the performance of agents on various programming tasks, focusing on aspects such as correctness and error handling.
Papers using ToolBench (10)
- Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection LearningCoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool RetrievalAdarubric: Task-adaptive Rubrics For LLM Agent EvaluationHow Many Tools Should an LLM Agent See? A Chance-Corrected AnswerCase-Based Calibration of Adaptive Reasoning and Execution for LLM Tool UseAgenther: Hindsight Experience Replay For LLM Agent Trajectory RelabelingBeyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM AgentsThink-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded ReasoningNaviAgent: Graph-Driven Bilevel Planning for Scalable Tool OrchestrationToolplanner: A Tool Augmented LLM For Multi Granularity Instructions With Path Planning And Feedback