← all papers · overview

Invisible Safety Threat: Malicious Finetuning For LLM Via Steganography

Abstract

Understanding and addressing potential safety alignment risks in large language models (LLMs) is critical for ensuring their safe and trustworthy deployment. In this paper, we highlight an insidious safety threat: a compromised LLM can maintain a facade of proper safety alignment while covertly generating harmful content. To achieve this, we finetune the model to understand and apply a steganograp

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).