← all datasets

Berkeley Function Calling Leaderboard v-3

Emerging
9papers using it
2024first seen

The 'Berkeley Function Calling Leaderboard v-3' is a benchmark dataset that contains 200 tasks used to evaluate the performance of function-calling language agents, particularly in relation to the effects of chain-of-thought reasoning on their accuracy.

Papers using Berkeley Function Calling Leaderboard v-3 (9)

Berkeley Function Calling Leaderboard v-3 dataset β€” papers, benchmarks & downloads Β· AI Agents