← all papers · overview

Gauss-newton Unlearning For The LLM Era

Abstract

Standard large language model training can create models that produce outputs their trainer deems unacceptable in deployment. The probability of these outputs can be reduced using methods such as LLM unlearning. However, unlearning a set of data (called the forget set) can degrade model performance on other distributions where the trainer wants to retain the model's behavior. To improve this trade

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).