← all datasets

M-3ToolEval

Emerging
2papers using it
2025first seen

The 'M-3ToolEval' is a benchmark dataset used to evaluate the reliability of code-mode tool use in models by assessing their performance on tasks involving inter-tool contract compliance and execution feedback.

Papers using M-3ToolEval (2)

M-3ToolEval dataset β€” papers, benchmarks & downloads Β· AI Agents