← all papers · overview

Prediction-powered Ranking Of Large Language Models

Abstract

Large language models are often ranked according to their level of alignment with human preferences -- a model is better than other models if its outputs are more frequently preferred by humans. One of the popular ways to elicit human preferences utilizes pairwise comparisons between the outputs provided by different models to the same inputs. However, since gathering pairwise comparisons by human

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).