AIME
Emerging9papers using it
2025first seen
The AIME dataset/benchmark contains competition math problems and is used to evaluate the performance of multiagent systems in solving these problems.
Papers using AIME (9)
- Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RLWhen LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence SignalsKV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time ScalingScaling Multiagent Systems With Process RewardsScRPO: From Errors to InsightsHERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMsThinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy OptimizationSIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical ReasoningAgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need