← all papers · overview

Unveiling Scoring Processes: Dissecting The Differences Between Llms And Human Graders In Automatic Scoring

Abstract

Large language models (LLMs) have demonstrated strong potential in performing automatic scoring for constructed response assessments. While constructed responses graded by humans are usually based on given grading rubrics, the methods by which LLMs assign scores remain largely unclear. It is also uncertain how closely AI's scoring process mirrors that of humans or if it adheres to the same grading

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).