← all papers · overview

Fine-grained Detoxification Via Instance-level Prefixes For Large Language Models

Abstract

Impressive results have been achieved in natural language processing (NLP) tasks through the training of large language models (LLMs). However, these models occasionally produce toxic content such as insults, threats, and profanity in response to certain prompts, thereby constraining their practical utility. To tackle this issue, various finetuning-based and decoding-based approaches have been uti

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).