← all papers · overview

Alignment-weighted DPO: A Principled Reasoning Approach To Improve Safety Alignment

Abstract

Recent advances in alignment techniques such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), and Direct Preference Optimization (DPO) have improved the safety of large language models (LLMs). However, these LLMs remain vulnerable to jailbreak attacks that disguise harmful intent through indirect or deceptive phrasing. Using causal intervention, we empirically de

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).