← all papers · overview

Pico: Peer Review In Llms Based On The Consistency Optimization

Abstract

Existing large language models (LLMs) evaluation methods typically focus on testing the performance on some closed-environment and domain-specific benchmarks with human annotations. In this paper, we explore a novel unsupervised evaluation direction, utilizing peer-review mechanisms to measure LLMs automatically. In this setting, both open-source and closed-source LLMs lie in the same environment,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).