← all papers · overview

Purifying Generative Llms From Backdoors Without Prior Knowledge Or Clean Reference

Abstract

Backdoor attacks pose severe security threats to large language models (LLMs), where a model behaves normally under benign inputs but produces malicious outputs when a hidden trigger appears. Existing backdoor removal methods typically assume prior knowledge of triggers, access to a clean reference model, or rely on aggressive finetuning configurations, and are often limited to classification task

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).