GPQA-Diamond
Emerging6papers using it
2025first seen
The 'GPQA-Diamond' dataset/benchmark contains multi-agent debate scenarios used to evaluate the mechanisms of convergence in reasoning among language models, specifically focusing on distinguishing between genuine deliberation and social compliance.
Papers using GPQA-Diamond (4)
- Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM DebateATLAS: Agentic Test-time Learning-to-Allocate ScalingLearning to Communicate: Toward End-to-End Optimization of Multi-Agent Language SystemsTwo Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning