← all papers · overview

Examining The Robustness Of LLM Evaluation To The Distributional Assumptions Of Benchmarks

Abstract

Benchmarks have emerged as the central approach for evaluating Large Language Models (LLMs). The research community often relies on a model's average performance across the test prompts of a benchmark to evaluate the model's performance. This is consistent with the assumption that the test prompts within a benchmark represent a random sample from a real-world distribution of interest. We note that

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).