← all papers · overview

Sharpness-aware Minimization In Logit Space Efficiently Enhances Direct Preference Optimization

Abstract

Direct Preference Optimization (DPO) has emerged as a popular algorithm for aligning pretrained large language models with human preferences, owing to its simplicity and training stability. However, DPO suffers from the recently identified squeezing effect (also known as likelihood displacement), where the probability of preferred responses decreases unintentionally during training. To understand

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).