HarmBench
Emerging8papers using it
8,156HF downloads
51HF likes
2025first seen
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal Data: Dataset About In this dataset card, we only use the behavior prompts proposed in HarmBench. License MIT Citation If you fin
π€ Hugging Faceβ mit
Papers using HarmBench (8)
- Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple AttacksEvoDefense: Co-Evolving Black-Box Defense with Large Language ModelsD-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output RewritingMembrane: A Self-Evolving Contrastive Safety Memory for LLM Agent DefenseBeyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team AttacksQuantifying LLM Safety Degradation Under Repeated Attacks Using Survival AnalysisTASO: Jailbreak LLMs via Alternative Template and Suffix OptimizationSoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models