← all datasets

Claw-Eval

Emerging
2papers using it
2026first seen

Claw-Eval End-to-end transparent benchmark for AI agents acting in the real world. Paper | Leaderboard | Code Dataset Structure Splits Split Examples Description general 161 Core agent tasks across 24 categories (communication, finance, ops, productivity, etc.) multimodal 101 Multimodal agentic tasks requiring percepti

Papers using Claw-Eval (2)

Claw-Eval dataset β€” papers, benchmarks & downloads Β· AI Agents