← all papers · overview

Realistic Evaluation Of Toxicity In Large Language Models

Abstract

Large language models (LLMs) have become integral to our professional workflows and daily lives. Nevertheless, these machine companions of ours have a critical flaw: the huge amount of data which endows them with vast and diverse knowledge, also exposes them to the inevitable toxicity and bias. While most LLMs incorporate defense mechanisms to prevent the generation of harmful content, these safeg

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).