Big-Math
Emerging2papers using it
2026first seen
The 'Big-Math' dataset/benchmark contains a collection of mathematical problems used to evaluate the reasoning capabilities of language models through their performance on these problems and the measurement of answer disagreement.