← all papers · overview

A Study Of Backdoors In Instruction Fine-tuned Language Models

Abstract

Backdoor data poisoning, inserted within instruction examples used to fine-tune a foundation Large Language Model (LLM) for downstream tasks (\textit\{e.g.,\} sentiment prediction), is a serious security concern due to the evasive nature of such attacks. The poisoning is usually in the form of a (seemingly innocuous) trigger word or phrase inserted into a very small fraction of the fine-tuning sam

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).