Camouflaged Jailbreak Prompts
Emerging1papers using it
2025first seen
The 'Camouflaged Jailbreak Prompts' dataset contains 500 curated examples of prompts—400 harmful and 100 benign—used to evaluate the effectiveness of safety protocols in large language models against subtle adversarial attacks that embed malicious intent within seemingly benign language.