← all papers · overview

Safeneuron: Neuron-level Safety Alignment For Large Language Models

Abstract

Large language models (LLMs) and multimodal LLMs are typically safety-aligned before release to prevent harmful content generation. However, recent studies show that safety behaviors are concentrated in a small subset of parameters, making alignment brittle and easily bypassed through neuron-level attacks. Moreover, most existing alignment methods operate at the behavioral level, offering limited

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).