← all papers · overview

Criterion-referenceability Determines Llm-as-a-judge Validity Across Physics Assessment Formats

Abstract

As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking can be trusted is essential. We evaluate LLM-as-a-judge marking across three physics assessment formats - structured questions, written essays, and scientific plots - comparing GPT-5.2, Grok 4.1, Claude Opus 4.5, DeepSeek-V3.2, Gemini Pro 3, and committee aggregations a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).