← all papers · overview

Xfinder: Large Language Models As Automated Evaluators For Reliable Evaluation

Abstract

The continuous advancement of large language models (LLMs) has brought increasing attention to the critical issue of developing fair and reliable methods for evaluating their performance. Particularly, the emergence of cheating phenomena, such as test set leakage and prompt format overfitting, poses significant challenges to the reliable evaluation of LLMs. As evaluation frameworks commonly use Re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).