← all papers · overview

Value Entanglement: Conflation Between Different Kinds Of Good In (some) Large Language Models

Abstract

Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Among the characteristics of value representation in humans is that they distinguish among value of different kinds. We investigate whether LLMs likewise distinguish three different kinds of good: moral, grammatical, and economic. By probing model behavior, emb

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).