← all papers · overview

Minor DPO Reject Penalty To Increase Training Robustness

Abstract

Learning from human preference is a paradigm used in large-scale language model (LLM) fine-tuning step to better align pretrained LLM to human preference for downstream task. In the past it uses reinforcement learning from human feedback (RLHF) algorithm to optimize the LLM policy to align with these preferences and not to draft too far from the original model. Recently, Direct Preference Optimiza

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).