← all papers · overview

Openeval: Benchmarking Chinese Llms Across Capability, Alignment And Safety

Abstract

The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluation platforms for assessing Chinese LLMs, many of these focus primarily on capabilities, usually overlooking potential alignment and safety issues. To address this gap, we introduce OpenEval, an evaluation testbed that b

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).