← all papers · overview

Quantifying Variance In Evaluation Benchmarks

Abstract

Evaluation benchmarks are the cornerstone of measuring capabilities of large language models (LLMs), as well as driving progress in said capabilities. Originally designed to make claims about capabilities (or lack thereof) in fully pretrained models, evaluation benchmarks are now also extensively used to decide between various training choices. Despite this widespread usage, we rarely quantify the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).