← all papers · overview

Wdpo: Winsorized Direct Preference Optimization For Robust LLM Alignment

Abstract

Direct Preference Optimization (DPO) aligns large language models by optimizing pairwise preferences and has shown remarkable effectiveness as a simple and scalable alternative to RLHF. However, in practice, preference data are often noisy. Existing robust variants of DPO mainly rely on uniform objective modifications or global reweighting. While partially effective, these methods treat noisy samp

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).