← all papers · overview

Memorization Or Interpolation ? Detecting LLM Memorization Through Input Perturbation Analysis

Abstract

While Large Language Models (LLMs) achieve remarkable performance through training on massive datasets, they can exhibit concerning behaviors such as verbatim reproduction of training data rather than true generalization. This memorization phenomenon raises significant concerns about data privacy, intellectual property rights, and the reliability of model evaluations. This paper introduces PEARL,