← all papers · overview

Criskeval: A Chinese Multi-level Risk Evaluation Benchmark Dataset For Large Language Models

Abstract

Large language models (LLMs) are possessed of numerous beneficial capabilities, yet their potential inclination harbors unpredictable risks that may materialize in the future. We hence propose CRiskEval, a Chinese dataset meticulously designed for gauging the risk proclivities inherent in LLMs such as resource acquisition and malicious coordination, as part of efforts for proactive preparedness. T

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).