← all papers · overview

Advbdgen: Adversarially Fortified Prompt-specific Fuzzy Backdoor Generator Against LLM Alignment

Abstract

With the growing adoption of reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs), the risk of backdoor installation during alignment has increased, leading to unintended and harmful behaviors. Existing backdoor triggers are typically limited to fixed word patterns, making them detectable during data cleaning and easily removable post-poisoning. In this work,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).