← all papers · overview

LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark For Chinese Large Language Models

Abstract

Chinese Large Language Models (LLMs) have recently demonstrated impressive capabilities across various NLP benchmarks and real-world applications. However, the existing benchmarks for comprehensively evaluating these LLMs are still insufficient, particularly in terms of measuring knowledge that LLMs capture. Current datasets collect questions from Chinese examinations across different subjects and

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).