← all papers · overview

Improving Few-shot Generalization Of Safety Classifiers Via Data Augmented Parameter-efficient Fine-tuning

Abstract

As large language models (LLMs) are widely adopted, new safety issues and policies emerge, to which existing safety classifiers do not generalize well. If we have only observed a few examples of violations of a new safety rule, how can we build a classifier to detect violations? In this paper, we study the novel setting of domain-generalized few-shot learning for LLM-based text safety classifiers.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).