← all papers · overview

NESTFUL: A Benchmark For Evaluating Llms On Nested Sequences Of API Calls

Abstract

The resurgence of autonomous agents built using large language models (LLMs) to solve complex real-world tasks has brought increased focus on LLMs' fundamental ability of tool or function calling. At the core of these agents, an LLM must plan, execute, and respond using external tools, APIs, and custom functions. Research on tool calling has gathered momentum, but evaluation benchmarks and dataset

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).