← all papers · overview

UNDIAL: Self-distillation With Adjusted Logits For Robust Unlearning In Large Language Models

Abstract

Mitigating the retention of sensitive or private information in large language models is essential for enhancing privacy and safety. Existing unlearning methods, like Gradient Ascent and Negative Preference Optimization, directly tune models to remove unwanted information. However, these methods often become unstable because they fine-tune by maximizing cross-entropy loss, which is the opposite of

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).