← all papers · overview

HPS: Hard Preference Sampling For Human Preference Alignment

Abstract

Aligning Large Language Model (LLM) responses with human preferences is vital for building safe and controllable AI systems. While preference optimization methods based on Plackett-Luce (PL) and Bradley-Terry (BT) models have shown promise, they face challenges such as poor handling of harmful content, inefficient use of dispreferred responses, and, specifically for PL, high computational costs. T

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).