← all papers · overview

3d-properties: Identifying Challenges In DPO And Charting A Path Forward

Abstract

Aligning large language models (LLMs) with human preferences has gained significant attention, with Proximal Policy Optimization (PPO) as a standard yet computationally expensive method and Direct Preference Optimization (DPO) as a more efficient alternative. While DPO offers simplicity, it remains underutilized in state-of-the-art LLMs, suggesting potential limitations. In this work, we revisit D

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).