Large Language Models (LLMs) have an unrivaled and invaluable ability to
"align" their output to a diverse range of human preferences, by mirroring them
in the text they generate. The internal characteristics of such models,
however, remain largely opaque. This work presents the Injectable Realignment
Model (IRM) as a novel approach to language model interpretability and
explainability. Inspired b
Related papers
Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).