← all papers · overview

Poisonbench: Assessing Large Language Model Vulnerability To Data Poisoning

Abstract

Preference learning is a central component for aligning current LLMs, but this process can be vulnerable to data poisoning attacks. To address this concern, we introduce PoisonBench, a benchmark for evaluating large language models' susceptibility to data poisoning during preference learning. Data poisoning attacks can manipulate large language model responses to include hidden malicious content o

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).