ACEBench
Emerging6papers using it
2025first seen
ACEBench Dataset This repository contains the ACEBench dataset, formatted for evaluating and training tool-using language models. The dataset has been processed into a unified structure, with problem descriptions merged with their corresponding ground-truth rubrics. Notebook used to format the dataset: Open in Colab Da
Papers using ACEBench (6)
- Towards General Agentic Intelligence via Environment ScalingScaling Agentic Capabilities via Grounded Interaction SynthesisMAVEN: Improving Generalization in Agentic Tool CallingControllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement LearningTry, Check and Retry: A Divide-and-Conquer Framework for Boosting Long-context Tool-Calling Performance of LLMsOn Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset