← all papers · overview

Soft Preference Optimization: Aligning Language Models To Expert Distributions

Abstract

We propose Soft Preference Optimization (SPO), a method for aligning generative models, such as Large Language Models (LLMs), with human preferences, without the need for a reward model. SPO optimizes model outputs directly over a preference dataset through a natural loss function that integrates preference loss with a regularization term across the model's entire output distribution rather than l

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).