← all papers · overview

Llms Do Not Grade Essays Like Humans

Abstract

Large language models have recently been proposed as tools for automated essay scoring, but their agreement with human grading remains unclear. In this work, we evaluate how LLM-generated scores compare with human grades and analyze the grading behavior of several models from the GPT and Llama families in an out-of-the-box setting, without task-specific training. Our results show that agreement be

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).