← all papers · overview

LLM Essay Scoring Under Holistic And Analytic Rubrics: Prompt Effects And Bias

Abstract

Despite growing interest in using Large Language Models (LLMs) for educational assessment, it remains unclear how closely they align with human scoring. We present a systematic evaluation of instruction-tuned LLMs across three open essay-scoring datasets (ASAP 2.0, ELLIPSE, and DREsS) that cover both holistic and analytic scoring. We analyze agreement with human consensus scores, directional bias,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).