9 utility benchmarks
Emerging1papers using it
2025first seen
The '9 utility benchmarks' dataset evaluates general knowledge, instruction following, and agentic workflows in language models.
The '9 utility benchmarks' dataset evaluates general knowledge, instruction following, and agentic workflows in language models.