← all datasets

GPQA-Diamond

Emerging
6papers using it
2025first seen

The 'GPQA-Diamond' dataset/benchmark contains multi-agent debate scenarios used to evaluate the mechanisms of convergence in reasoning among language models, specifically focusing on distinguishing between genuine deliberation and social compliance.

Papers using GPQA-Diamond (4)

GPQA-Diamond dataset β€” papers, benchmarks & downloads Β· AI Agents