← all papers · overview

Stop Overvaluing Multi-agent Debate -- We Must Rethink Evaluation And Embrace Model Heterogeneity

Abstract

Multi-agent debate (MAD) has gained significant attention as a promising line of research to improve the factual accuracy and reasoning capabilities of large language models (LLMs). Despite its conceptual appeal, current MAD research suffers from critical limitations in evaluation practices, including limited benchmark coverage, weak baseline comparisons, and inconsistent setups. This paper presen

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).