← all papers · overview

Rubriceval: A Rubric-level Meta-evaluation Benchmark For LLM Judges In Instruction Following

Abstract

Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these rubric-level evaluations remains unclear, calling for meta-evaluation. However, prior meta-evaluation efforts largely focus on the response level, failing to assess the fine-grained judgment accuracy that rubric-based ev

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).