← all datasets

AQuA

Emerging
2papers using it
2025first seen

The 'AQUA' dataset/benchmark is used to evaluate the mathematical reasoning capabilities of large language models (LLMs) by providing a set of problems that require formal verification through theorem provers.

Papers using AQuA (2)

AQuA dataset β€” papers, benchmarks & downloads Β· Reinforcement Learning