← all papers · overview

Prompt-to-leaderboard

Abstract

Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures user- and prompt-specific variations in model performance. To address this, we propose Prompt-to-Leaderboard (P2L), a method that produces leaderboards specific to a prompt. The core idea is to train an LLM taking natural languag

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).