← all papers · overview

How Reliable Are Automatic Evaluation Methods For Instruction-tuned Llms?

Abstract

Work on instruction-tuned Large Language Models (LLMs) has used automatic methods based on text overlap and LLM judgments as cost-effective alternatives to human evaluation. In this paper, we perform a meta-evaluation of such methods and assess their reliability across a broad range of tasks. In evaluating how well automatic methods align with human evaluations, correlation metrics are the most co

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).