← all papers · overview

Dissecting Human And LLM Preferences

Abstract

As a relative quality comparison of model responses, human and Large Language Model (LLM) preferences serve as common alignment goals in model fine-tuning and criteria in evaluation. Yet, these preferences merely reflect broad tendencies, resulting in less explainable and controllable models with potential safety risks. In this work, we dissect the preferences of human and 32 different LLMs to und

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).