Math-Hard
Emerging1papers using it
2025first seen
The 'Math-Hard' dataset is a benchmark used to evaluate the mathematical reasoning capabilities of Large Language Models (LLMs).
The 'Math-Hard' dataset is a benchmark used to evaluate the mathematical reasoning capabilities of Large Language Models (LLMs).