This paper presents reports on a series of experiments with a novel dataset
evaluating how well Large Language Models (LLMs) can mark (i.e. grade) open
text responses to short answer questions, Specifically, we explore how well
different combinations of GPT version and prompt engineering strategies
performed at marking real student answers to short answer across different
domain areas (Science and
Related papers
Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).