AQuA
Emerging2papers using it
2025first seen
The 'AQUA' dataset/benchmark is used to evaluate the mathematical reasoning capabilities of large language models (LLMs) by providing a set of problems that require formal verification through theorem provers.