GPQA
Emerging20papers using it
99,381HF downloads
492HF likes
2024first seen
Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with ful
π€ Hugging Faceβ cc-by-4.0
Papers using GPQA (20)
- Stable Reinforcement Learning for Efficient ReasoningRUMAD: Reinforcement-Unifying Multi-Agent DebateRL with Learnable Textual Feedback: A Bilevel ApproachLearning to Orchestrate Agents in Natural Language with the ConductorAdaptive Graph of Thoughts: Test-Time Adaptive Reasoning Unifying Chain,
Tree, and Graph StructuresApriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient ReasoningOff-Policy Value-Based Reinforcement Learning for Large Language ModelsCLEANER: Self-Purified Trajectories Boost Agentic Reinforcement LearningCan LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM ReasoningOnline Rubrics Elicitation from Pairwise ComparisonsReasoning with Sampling: Your Base Model is Smarter Than You ThinkTrain for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates HallucinationsSample More to Think Less: Group Filtered Policy Optimization for Concise ReasoningRL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement LearningSHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement LearningEnigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable PuzzlesReinforcing General Reasoning without VerifiersMaximizing Confidence Alone Improves ReasoningSynthetic Data RL: Task Definition Is All You NeedCPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning
Tasks