← all papers · overview

Backdooring Bias In Large Language Models

Abstract

Large language models (LLMs) are increasingly deployed in settings where inducing a bias toward a certain topic can have significant consequences, and backdoor attacks can be used to produce such models. Prior work on backdoor attacks has largely focused on a black-box threat model, with an adversary targeting the model builder's LLM. However, in the bias manipulation setting, the model builder th

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).