Benchmarks
loadingβ¦
loadingβ¦
Benchmarks is one of the most active areas in Awesome AI Agents β 2,602 papers in this collection, evaluated on datasets like ALFWorld, GAIA, LoCoMo. A strong starting point is "R-Zero: Self-Evolving Reasoning LLM from Zero Data".