AMC-23
Emerging13papers using it
2024first seen
The 'AMC-23' dataset/benchmark is used to evaluate the performance of large language models in reasoning tasks.
Papers using AMC-23 (13)
- Nemotron-CrossThink: Scaling Self-Learning beyond Math ReasoningSqueeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language ModelPopuLoRA: Co-Evolving LLM Populations for Reasoning Self-PlayLearning-Zone Energy: Online Data Selection for Efficient RL Post-TrainingThink Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient ReasoningBeyond Variance: Prompt-Efficient RLVR via Rare-Event Amplification and Bidirectional PairingLong Chain-of-Thought Compression via Fine-Grained Group Policy OptimizationMasked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable RewardsConfidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative DraftsStop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for ReasoningReinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn'tQwen2.5-Math Technical Report: Toward Mathematical Expert Model via
Self-Improvement