← all papers · overview

Strong Preferences Affect The Robustness Of Preference Models And Value Alignment

Abstract

Value alignment, which aims to ensure that large language models (LLMs) and other AI agents behave in accordance with human values, is critical for ensuring safety and trustworthiness of these systems. A key component of value alignment is the modeling of human preferences as a representation of human values. In this paper, we investigate the robustness of value alignment by examining the sensitiv

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).