← all papers · overview

Learning And Forgetting Unsafe Examples In Large Language Models

Abstract

As the number of large language models (LLMs) released to the public grows, there is a pressing need to understand the safety implications associated with these models learning from third-party custom finetuning data. We explore the behavior of LLMs finetuned on noisy custom data containing unsafe content, represented by datasets that contain biases, toxicity, and harmfulness, finding that while a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).