← all papers · overview

The Mysterious Case Of Neuron 1512: Injectable Realignment Architectures Reveal Internal Characteristics Of Meta's Llama 2 Model

Abstract

Large Language Models (LLMs) have an unrivaled and invaluable ability to "align" their output to a diverse range of human preferences, by mirroring them in the text they generate. The internal characteristics of such models, however, remain largely opaque. This work presents the Injectable Realignment Model (IRM) as a novel approach to language model interpretability and explainability. Inspired b

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).