← all papers · overview

Rocketeval: Efficient Automated LLM Evaluation Via Grading Checklist

Abstract

Evaluating large language models (LLMs) in diverse and challenging scenarios is essential to align them with human preferences. To mitigate the prohibitive costs associated with human evaluations, utilizing a powerful LLM as a judge has emerged as a favored approach. Nevertheless, this methodology encounters several challenges, including substantial expenses, concerns regarding privacy and securit

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).