← all papers · overview

Making Harmful Behaviors Unlearnable For Large Language Models

Abstract

Large language models (LLMs) have shown great potential as general-purpose AI assistants in various domains. To meet the requirements of different applications, LLMs are often customized by further fine-tuning. However, the powerful learning ability of LLMs not only enables them to acquire new tasks but also makes them susceptible to learning undesired behaviors. For example, even safety-aligned L

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).