AdvBench
Emerging13papers using it
16,127HF downloads
110HF likes
2025first seen
Dataset Card for AdvBench Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models Data: AdvBench Dataset About AdvBench is a set of 500 harmful behaviors formulated as instructions. These behaviors range over the same themes as the harmful strings setting, but the adversary’s goal is instead to
Hugging Face⚖ mit
Papers using AdvBench (13)
- Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple AttacksEvoDefense: Co-Evolving Black-Box Defense with Large Language ModelsHidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful BehaviorJailbreak Foundry: From Papers to Runnable Attacks for Reproducible BenchmarkingMaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language ModelsEvolving Skill-Structured Attack Memory Enhances LLM JailbreakingRefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMsSafety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic AnalysisRefusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)TASO: Jailbreak LLMs via Alternative Template and Suffix OptimizationJailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language ModelsCon Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual ModalitiesGCG Attack On A Diffusion LLM