← all papers · overview

Noisy But Valid: Robust Statistical Evaluation Of Llms With Imperfect Judges

Abstract

Reliable certification of Large Language Models (LLMs)-verifying that failure rates are below a safety threshold-is critical yet challenging. While "LLM-as-a-Judge" offers scalability, judge imperfections, noise, and bias can invalidate statistical guarantees. We introduce a "Noisy but Valid" hypothesis testing framework to address this. By leveraging a small human-labelled calibration set to esti

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).