← all papers · overview

Intrinsic Self-correction For Enhanced Morality: An Analysis Of Internal Mechanisms And The Superficial Hypothesis

Abstract

Large Language Models (LLMs) are capable of producing content that perpetuates stereotypes, discrimination, and toxicity. The recently proposed moral self-correction is a computationally efficient method for reducing harmful content in the responses of LLMs. However, the process of how injecting self-correction instructions can modify the behavior of LLMs remains under-explored. In this paper, we

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).