← all datasets

BFCL

Emerging
18papers using it
2026first seen

The BFCL dataset/benchmark contains instances of tool-selection failures in LLM agents and is used to evaluate the attention mechanisms and decision-making processes of these models in selecting the correct tools.

Papers using BFCL (18)

BFCL dataset β€” papers, benchmarks & downloads Β· AI Agents