← all papers · overview

Neurel-attack: Neuron Relearning For Safety Disalignment In Large Language Models

Abstract

Safety alignment in large language models (LLMs) is achieved through fine-tuning mechanisms that regulate neuron activations to suppress harmful content. In this work, we propose a novel approach to induce disalignment by identifying and modifying the neurons responsible for safety constraints. Our method consists of three key steps: Neuron Activation Analysis, where we examine activation patterns

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).