← all papers · overview

Rethinking Generative Large Language Model Evaluation For Semantic Comprehension

Abstract

Despite their sophisticated capabilities, large language models (LLMs) encounter a major hurdle in effective assessment. This paper first revisits the prevalent evaluation method-multiple choice question answering (MCQA), which allows for straightforward accuracy measurement. Through a comprehensive evaluation of 24 models across 11 benchmarks, we highlight several potential drawbacks of MCQA, for

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).