← all papers · overview

DEP: A Decentralized Large Language Model Evaluation Protocol

Abstract

With the rapid development of Large Language Models (LLMs), a large number of benchmarks have been proposed. However, most benchmarks lack unified evaluation standard and require the manual implementation of custom scripts, making results hard to ensure consistency and reproducibility. Furthermore, mainstream evaluation frameworks are centralized, with datasets and answers, which increases the ris

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).