← all papers · overview

Explainable LLM Unlearning Through Reasoning

Abstract

LLM unlearning is essential for mitigating safety, copyright, and privacy concerns in pre-trained large language models (LLMs). Compared to preference alignment, it offers a more explicit way by removing undesirable knowledge characterized by specific unlearning datasets. In previous works, gradient ascent (GA) and its variants have shown promise for implementing unlearning, yet their untargeted n

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).