← all papers · overview

Resurrecting saturated LLM benchmarks with adversarial encoding

Abstract

Recent work showed that small changes in benchmark questions can reduce LLMs' reasoning and recall. We explore two such changes: pairing questions and adding more answer options, on three benchmarks: WMDP-bio, GPQA, and MMLU variants. We find that for more capable models, these predictably reduce performance, essentially heightening the performance ceiling of a benchmark and unsaturating it again. We suggest this approach can resurrect old benchmarks.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).