13 challenging benchmarks
Emerging1papers using it
2025first seen
The '13 challenging benchmarks' dataset consists of a variety of tasks designed to evaluate the performance of large language models (LLMs) in single-turn reasoning scenarios and their ability to interact with external tools in multi-turn contexts.