← all papers · overview

PRE: A Peer Review Based Large Language Model Evaluator

Abstract

The impressive performance of large language models (LLMs) has attracted considerable attention from the academic and industrial communities. Besides how to construct and train LLMs, how to effectively evaluate and compare the capacity of LLMs has also been well recognized as an important yet difficult problem. Existing paradigms rely on either human annotators or model-based evaluators to evaluat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).