← all papers · overview

Generalization Or Memorization: Data Contamination And Trustworthy Evaluation For Large Language Models

Abstract

Recent statements about the impressive capabilities of large language models (LLMs) are usually supported by evaluating on open-access benchmarks. Considering the vast size and wide-ranging sources of LLMs' training data, it could explicitly or implicitly include test data, leading to LLMs being more susceptible to data contamination. However, due to the opacity of training data, the black-box acc

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).