AIME-25
Emerging26papers using it
2025first seen
AIME-25 is a benchmark dataset used to evaluate reinforcement learning with verifiable rewards (RLVR) in the context of solving challenging math questions.
Papers using AIME-25 (26)
- QuestA: Expanding Reasoning Capacity in LLMs via Question AugmentationRecursive Self-Aggregation Unlocks Deep Thinking in Large Language ModelsLearn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningBeyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningHTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective ControlLearning-Zone Energy: Online Data Selection for Efficient RL Post-TrainingFrom Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM ReasoningMitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint AnnealingLearn Hard Problems During RL with Reference Guided Fine-tuningThink Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient ReasoningiGRPO: Self-Feedback-Driven LLM ReasoningLatent Poincar\'e Shaping for Agentic Reinforcement LearningDataChef: Cooking Up Optimal Data Recipes for LLM Adaptation via Reinforcement LearningTransformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language ModelsCan LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM ReasoningShorter but not Worse: Frugal Reasoning via Easy Samples as Length Regularizers in Math RLVRMasked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable RewardsLearning on the Job: Test-Time Curricula for Targeted Reinforcement LearningA$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid ReasoningTowards High Data Efficiency in Reinforcement Learning with Verifiable RewardDCPO: Dynamic Clipping Policy OptimizationSingle-stream Policy OptimizationEvolving Language Models without Labels: Majority Drives Selection, Novelty Promotes VariationPromoting Efficient Reasoning with Verifiable Stepwise RewardOn the Design of KL-Regularized Policy Gradient Algorithms for LLM ReasoningSkywork Open Reasoner 1 Technical Report