← all papers · overview

Pacost: Paired Confidence Significance Testing For Benchmark Contamination Detection In Large Language Models

Abstract

Large language models (LLMs) are known to be trained on vast amounts of data, which may unintentionally or intentionally include data from commonly used benchmarks. This inclusion can lead to cheatingly high scores on model leaderboards, yet result in disappointing performance in real-world applications. To address this benchmark contamination problem, we first propose a set of requirements that p

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).