← all papers · overview

Simplesafetytests: A Test Suite For Identifying Critical Safety Risks In Large Language Models

Abstract

The past year has seen rapid acceleration in the development of large language models (LLMs). However, without proper steering and safeguards, LLMs will readily follow malicious instructions, provide unsafe advice, and generate toxic content. We introduce SimpleSafetyTests (SST) as a new test suite for rapidly and systematically identifying such critical safety risks. The test suite comprises 100

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).