← all papers · overview

Locking Down The Finetuned Llms Safety

Abstract

Fine-tuning large language models (LLMs) on additional datasets is often necessary to optimize them for specific downstream tasks. However, existing safety alignment measures, which restrict harmful behavior during inference, are insufficient to mitigate safety risks during fine-tuning. Alarmingly, fine-tuning with just 10 toxic sentences can make models comply with harmful instructions. We introd

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).