← all papers · overview

PATCH! psychometrics-assisted Benchmarking Of Large Language Models Against Human Populations: A Case Study Of Proficiency In 8th Grade Mathematics

Abstract

Many existing benchmarks of large (multimodal) language models (LLMs) focus on measuring LLMs' academic proficiency, often with also an interest in comparing model performance with human test takers'. While such benchmarks have proven key to the development of LLMs, they suffer from several limitations, including questionable measurement quality (e.g., Do they measure what they are supposed to in

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).