AGIEval-en
Emerging2papers using it
2025first seen
AGIEval-en is a benchmark used to evaluate the complex reasoning abilities of large language models (LLMs) through reasoning-intensive tasks.
AGIEval-en is a benchmark used to evaluate the complex reasoning abilities of large language models (LLMs) through reasoning-intensive tasks.