AIME-25
Emerging7papers using it
2025first seen
'AIME'25' is a benchmark dataset that contains math problems used to evaluate the effectiveness of system prompt optimization methods in improving agent performance on mathematical tasks.
Papers using AIME-25 (7)
- Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and
BeyondSePO: Self-Evolving Prompt Agent for System Prompt OptimizationZAYA1-8B Technical ReportCut Your Losses! Learning to Prune Paths Early for Efficient Parallel ReasoningPromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model ReasoningrStar2-Agent: Agentic Reasoning Technical ReportPromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model
Reasoning