← all papers · overview

No Free Labels: Limitations Of Llm-as-a-judge Without Human Grounding

Abstract

Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate response quality, is appealing due to its scalability, low cost, and strong correlations with human stylistic preferences. However, it remains unclear how accurately

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).