← all papers · overview

When The Model Said 'no Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, And Safety Was Terrified

Abstract

Large Language Models (LLMs) need to be in accordance with human values-being helpful, harmless, and honest (HHH)-is important for safe deployment. Existing works use Supervised Fine-Tuning (SFT) and Mixture-of-Experts (MoE) to align LLMs. However, these works face challenges in multi-objective settings, such as SFT leading to interference between conflicting objectives, while MoEs suffer from mis

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).