← all papers · overview

Nonparametric LLM Evaluation From Preference Data

Abstract

Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards. However, many existing approaches either rely on restrictive parametric assumptions or lack valid uncertainty quantification when flexible machine learning methods are used. In this paper, we propose a nonparametric statistical framework, DMLEval, for comparing and ranki

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).