← all papers · overview

Toolscan: A Benchmark For Characterizing Errors In Tool-use Llms

Abstract

Evaluating Large Language Models (LLMs) is one of the most critical aspects of building a performant compound AI system. Since the output from LLMs propagate to downstream steps, identifying LLM errors is crucial to system performance. A common task for LLMs in AI systems is tool use. While there are several benchmark environments for evaluating LLMs on this task, they typically only give a succes

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).