← all papers · overview

Systematic Characterization Of The Effectiveness Of Alignment In Large Language Models For Categorical Decisions

Abstract

As large language models (LLMs) are deployed in high-stakes domains like healthcare, understanding how well their decision-making aligns with human preferences and values becomes crucial, especially when we recognize that there is no single gold standard for these preferences. This paper applies a systematic methodology for evaluating preference alignment in LLMs on categorical decision-making wit

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).