← all papers · overview

Are Language Models Sensitive To Morally Irrelevant Distractors?

Abstract

With the rapid development and uptake of large language models (LLMs) across high-stakes settings, it is increasingly important to ensure that LLMs behave in ways that align with human values. Existing moral benchmarks prompt LLMs with value statements, moral scenarios, or psychological questionnaires, with the implicit underlying assumption that LLMs report somewhat stable moral preferences. Howe

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).