← all papers · overview

Safety Alignment Of Large Language Models Via Contrasting Safe And Harmful Distributions

Abstract

With the widespread application of Large Language Models (LLMs), it has become a significant concern to ensure their safety and prevent harmful responses. While current safe-alignment methods based on instruction fine-tuning and Reinforcement Learning from Human Feedback (RLHF) can effectively reduce harmful responses from LLMs, they often require high-quality datasets and heavy computational over

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).