← all datasets

AIME-24

Emerging
6papers using it
2025first seen

The 'AIME 24' dataset/benchmark contains a collection of problems designed to evaluate the reasoning capabilities of large language models (LLMs) in tasks such as mathematics and programming.

Papers using AIME-24 (6)

AIME-24 dataset β€” papers, benchmarks & downloads Β· AI for Code