MATH-500
Emerging15papers using it
2025first seen
Dataset Card for MATH-500 This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
Papers using MATH-500 (15)
- Geometric Signatures of Reasoning: A Spectral Perspective on Task HardnessSPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language ModelsTreeThink: A Modular Tree Search Library for Mathematical Reasoning with LLMsSLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory SlicingMaximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement LearningSAGE-32B: Agentic Reasoning Via Iterative DistillationCounterfactual Credit Policy Optimization for Multi-Agent CollaborationWhat If We Allocate Test-Time Compute Adaptively?PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language ModelsScRPO: From Errors to InsightsSIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical ReasoningFrom Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMsReinforce LLM Reasoning through Multi-Agent ReflectionA*-Decoding: Token-Efficient Inference ScalingWalk Before You Run! Concise LLM Reasoning via Reinforcement Learning