← all papers · overview

Benchmark Data Contamination Of Large Language Models: A Survey

Abstract

The rapid development of Large Language Models (LLMs) like GPT-4, Claude-3, and Gemini has transformed the field of natural language processing. However, it has also resulted in a significant issue known as Benchmark Data Contamination (BDC). This occurs when language models inadvertently incorporate evaluation benchmark information from their training data, leading to inaccurate or unreliable per

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).