Anthropic HH
Emerging2papers using it
2024first seen
The 'ANTHROPIC-HH' dataset is a benchmark that contains prompts reclassified into safety domains to evaluate the pluralistic sensitivity of large language models in reflecting diverse moral norms and cultural expectations across different demographic contexts.