StrongREJECT
Emerging2papers using it
4,038HF downloads
23HF likes
2025first seen
StrongREJECT A novel benchmark of 313 malicious prompts for use in evaluating jailbreaking attacks against LLMs, aimed to expose whether a jailbreak attack actually enables malicious actors to utilize LLMs for harmful tasks. Dataset link: https://github.com/alexandrasouly/strongreject/blob/main/strongreject_dataset/str
π€ Hugging Faceβ mit