← all papers · overview

Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective

Abstract

Direct Preference Optimization (DPO) has gained attention as an efficient alternative to reinforcement learning from human feedback (RLHF) for aligning large language models (LLMs) with human preferences. Despite its advantages, DPO suffers from a length bias, generating responses longer than those from the reference model. Existing solutions like SimPO and SamPO address this issue but uniformly t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).