← all datasets

Arena-Hard

Emerging
13papers using it
84HF downloads
1HF likes
2024first seen

The 'Arena-Hard' dataset is a benchmark used to evaluate the performance of LLMs in alignment tasks by providing challenging scenarios that require reasoning and decision-making without verifiable ground-truth verifiers.

Papers using Arena-Hard (13)

Arena-Hard dataset β€” papers, benchmarks & downloads Β· Reinforcement Learning