← all papers · overview

Rethinking LLM Memorization Through The Lens Of Adversarial Compression

Abstract

Large language models (LLMs) trained on web-scale datasets raise substantial concerns regarding permissible data usage. One major question is whether these models "memorize" all their training data or they integrate many data sources in some way more akin to how a human would learn and synthesize information. The answer hinges, to a large degree, on how we define memorization. In this work, we pro

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).