← all papers · overview

Training On The Benchmark Is Not All You Need

Abstract

The success of Large Language Models (LLMs) relies heavily on the huge amount of pre-training data learned in the pre-training phase. The opacity of the pre-training process and the training data causes the results of many benchmark tests to become unreliable. If any model has been trained on a benchmark test set, it can seriously hinder the health of the field. In order to automate and efficientl

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).