AIME-24
Emerging6papers using it
2025first seen
The 'AIME 24' dataset/benchmark contains a collection of problems designed to evaluate the reasoning capabilities of large language models (LLMs) in tasks such as mathematics and programming.
Papers using AIME-24 (6)
- Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and
BeyondPromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model ReasoningrStar2-Agent: Agentic Reasoning Technical ReportPromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model
ReasoningWhich Data Attributes Stimulate Math and Code Reasoning? An
Investigation via Influence FunctionsWhich Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions