← all papers · overview

When Can We Trust LLM Graders? Calibrating Confidence For Automated Assessment

Abstract

Large Language Models (LLMs) show promise for automated grading, but their outputs can be unreliable. Rather than improving grading accuracy directly, we address a complementary problem: \textit\{predicting when an LLM grader is likely to be correct\}. This enables selective automation where high-confidence predictions are processed automatically while uncertain cases are flagged for human review.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).