← all papers · overview

Acebench: Who Wins The Match Point In Tool Usage?

Abstract

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex problems. However, existing benchmarks for evaluating LLMs' tool usage face several limitations: (1) limited evaluation scenarios, often lacking assessments in real multi-turn dialogue contexts; (2) narrow evaluation dim

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).