← all papers · overview

Silent Sabotage During Fine-tuning: Few-shot Rationale Poisoning Of Compact Medical Llms

Abstract

Supervised fine-tuning (SFT) is essential for the development of medical large language models (LLMs), yet prior poisoning studies have mainly focused on the detectable backdoor attacks. We propose a novel poisoning attack targeting the reasoning process of medical LLMs during SFT. Unlike backdoor attacks, our method injects poisoned rationales into few-shot training data, leading to stealthy degr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).