← all papers · overview

Beyond Static Evaluation: A Dynamic Approach To Assessing AI Assistants' API Invocation Capabilities

Abstract

With the rise of Large Language Models (LLMs), AI assistants' ability to utilize tools, especially through API calls, has advanced notably. This progress has necessitated more accurate evaluation methods. Many existing studies adopt static evaluation, where they assess AI assistants' API call based on pre-defined dialogue histories. However, such evaluation method can be misleading, as an AI assis

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).