← all papers · overview

Online DPO: Online Direct Preference Optimization With Fast-slow Chasing

Abstract

Direct Preference Optimization (DPO) improves the alignment of large language models (LLMs) with human values by training directly on human preference datasets, eliminating the need for reward models. However, due to the presence of cross-domain human preferences, direct continual training can lead to catastrophic forgetting, limiting DPO's performance and efficiency. Inspired by intraspecific com

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).