← all papers · overview

LSSF: Safety Alignment For Large Language Models Through Low-rank Safety Subspace Fusion

Abstract

The safety mechanisms of large language models (LLMs) exhibit notable fragility, as even fine-tuning on datasets without harmful content may still undermine their safety capabilities. Meanwhile, existing safety alignment methods predominantly rely on the fine-tuning process, which inadvertently leads to the increased complexity and computational resources required. To address these issues, we intr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).