← all papers · overview

Representation Noising: A Defence Mechanism Against Harmful Finetuning

Abstract

Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs). While safety measures like preventing jailbreaks and improving safety guardrails are important, such me

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).