← all papers · overview

Mechanistic Origin Of Moral Indifference In Language Models

Abstract

Existing behavioral alignment techniques for Large Language Models (LLMs) often neglect the discrepancy between surface compliance and internal unaligned representations, leaving LLMs vulnerable to long-tail risks. More crucially, we posit that LLMs possess an inherent state of moral indifference due to compressing distinct moral concepts into uniform probability distributions. We verify and remed

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).