← all papers · overview

Metabench -- A Sparse Benchmark Of Reasoning And Knowledge In Large Language Models

Abstract

Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the Open LLM Leaderboard aim to quantify these differences with several large benchmarks (sets of test items to which an LLM can respond either correctly or incorrectly). However, high correlations within and between benchmark scores suggest that (1) there exists a small set of common underlying abilities

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).