← all papers · overview

Model Editing As A Robust And Denoised Variant Of DPO: A Case Study On Toxicity

Abstract

Recent alignment algorithms such as direct preference optimization (DPO) have been developed to improve the safety of large language models (LLMs) by training these models to match human behaviors exemplified by preference data. However, these methods are both computationally intensive and lacking in controllability and transparency, inhibiting their widespread use. Furthermore, these tuning-based

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).