← all papers · overview

Compositional Preference Models For Aligning Lms

Abstract

As language models (LMs) become more capable, it is increasingly important to align them with human preferences. However, the dominant paradigm for training Preference Models (PMs) for that purpose suffers from fundamental limitations, such as lack of transparency and scalability, along with susceptibility to overfitting the preference dataset. We propose Compositional Preference Models (CPMs), a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).