← all datasets

AIME-24/25

Emerging
4papers using it
2025first seen

The 'AIME-24/25' dataset/benchmark contains a collection of tasks designed to evaluate the performance and robustness of reinforcement learning algorithms, particularly in the context of agentic problem-solving with Large Language Models.

Papers using AIME-24/25 (4)

AIME-24/25 dataset β€” papers, benchmarks & downloads Β· Reinforcement Learning