← all papers · overview

Pre-dpo: Improving Data Utilization In Direct Preference Optimization Using A Guiding Reference Model

Abstract

Direct Preference Optimization (DPO) simplifies reinforcement learning from human feedback (RLHF) for large language models (LLMs) by directly optimizing human preferences without an explicit reward model. We find that during DPO training, the reference model plays the role of a data weight adjuster. However, the common practice of initializing the policy and reference models identically in DPO ca

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).