← all papers · overview

To Each (textual Sequence) Its Own: Improving Memorized-data Unlearning In Large Language Models

Abstract

LLMs have been found to memorize training textual sequences and regurgitate verbatim said sequences during text generation time. This fact is known to be the cause of privacy and related (e.g., copyright) problems. Unlearning in LLMs then takes the form of devising new algorithms that will properly deal with these side-effects of memorized data, while not hurting the model's utility. We offer a fr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).