Arena-Hard
Emerging13papers using it
84HF downloads
1HF likes
2024first seen
The 'Arena-Hard' dataset is a benchmark used to evaluate the performance of LLMs in alignment tasks by providing challenging scenarios that require reasoning and decision-making without verifiable ground-truth verifiers.
Papers using Arena-Hard (13)
- QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable RewardsReferences Improve LLM Alignment in Non-Verifiable DomainsReward Model Routing in AlignmentOnline Rubrics Elicitation from Pairwise ComparisonsTGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference OptimizationPretrain Value, Not Reward: Decoupled Value Policy OptimizationScalable Reinforcement Post-Training Beyond Static Human Prompts:
Evolving Alignment via Asymmetric Self-PlayDPO Meets PPO: Reinforced Token Optimization for RLHFSimPO: Simple Preference Optimization with a Reference-Free RewardRLHF Workflow: From Reward Modeling to Online RLHFThe Perfect Blend: Redefining RLHF with Mixture of JudgesAlphaDPO: Adaptive Reward Margin for Direct Preference OptimizationT-REG: Preference Optimization with Token-Level Reward Regularization