← all papers · overview

Are Large Language Models Good Essay Graders?

Abstract

We evaluate the effectiveness of Large Language Models (LLMs) in assessing essay quality, focusing on their alignment with human grading. More precisely, we evaluate ChatGPT and Llama in the Automated Essay Scoring (AES) task, a crucial natural language processing (NLP) application in Education. We consider both zero-shot and few-shot learning and different prompting approaches. We compare the num

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).