← all datasets

AIME

Emerging
33papers using it
171HF downloads
0HF likes
2025first seen

The AIME dataset/benchmark contains mathematical reasoning tasks used to evaluate the performance of large language models in generating correct solutions and intermediate reasoning steps.

Papers using AIME (33)

AIME dataset β€” papers, benchmarks & downloads Β· Reinforcement Learning