← all papers · overview

Measuring Taiwanese Mandarin Language Understanding

Abstract

The evaluation of large language models (LLMs) has drawn substantial attention in the field recently. This work focuses on evaluating LLMs in a Chinese context, specifically, for Traditional Chinese which has been largely underrepresented in existing benchmarks. We present TMLU, a holistic evaluation suit tailored for assessing the advanced knowledge and reasoning capability in LLMs, under the con

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).