← all papers · overview

Guardrail Baselines For Unlearning In Llms

Abstract

Recent work has demonstrated that finetuning is a promising approach to 'unlearn' concepts from large language models. However, finetuning can be expensive, as it requires both generating a set of examples and running iterations of finetuning to update the model. In this work, we show that simple guardrail-based approaches such as prompting and filtering can achieve unlearning results comparable t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).