← all papers · overview

Same Input, Different Scores: A Multi Model Study On The Inconsistency Of LLM Judge

Fiona Lau·2026

Abstract

Large language models are increasingly used as automated evaluators in research and enterprise settings, a practice known as LLM-as-a-judge. While prior work has examined accuracy, bias, and alignment with human preferences, far less attention has been given to how consistently LLMs assign numerical scores, an important concern for many production workflows. This study systematically evaluates sco

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).