OlympiadBench
Emerging15papers using it
1,462HF downloads
5HF likes
2025first seen
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems[ACL 2024] π arXiv | GitHub Note: We have made adjustments to the image content in the multimodal portion of the dataset and fixed previous issues where some images in the English physics subset were no
Papers using OlympiadBench (15)
- Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language ModelCATPO: Critique-Augmented Tree Policy OptimizationMaximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement LearningPopuLoRA: Co-Evolving LLM Populations for Reasoning Self-PlaySAGE: Multi-Agent Self-Evolution for LLM ReasoningPAPO: Stabilizing Rubric Integration Training via Decoupled Advantage NormalizationASI-Evolve: AI Accelerates AIPrAg-PO: Prompt Augmented Policy Optimization for Robust and Diverse Mathematical ReasoningEMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KLEBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy OptimizationTransformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language ModelsWirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement LearningSPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative RolloutsConfidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts