← all papers · overview

Ranking Large Language Models Without Ground Truth

Abstract

Evaluation and ranking of large language models (LLMs) has become an important problem with the proliferation of these models and their impact. Evaluation methods either require human responses which are expensive to acquire or use pairs of LLMs to evaluate each other which can be unreliable. In this paper, we provide a novel perspective where, given a dataset of prompts (viz. questions, instructi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).