← all papers · overview

DESTEIN: Navigating Detoxification Of Language Models Via Universal Steering Pairs And Head-wise Activation Fusion

Abstract

Despite the remarkable achievements of language models (LMs) across a broad spectrum of tasks, their propensity for generating toxic outputs remains a prevalent concern. Current solutions involving finetuning or auxiliary models usually require extensive computational resources, hindering their practicality in large language models (LLMs). In this paper, we propose DeStein, a novel method that det

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).