← all papers · overview

Spark-llm-eval: A Distributed Framework For Statistically Rigorous Large Language Model Evaluation

Abstract

Evaluating large language models at scale remains a practical bottleneck for many organizations. While existing evaluation frameworks work well for thousands of examples, they struggle when datasets grow to hundreds of thousands or millions of samples. This scale is common when assessing model behavior across diverse domains or conducting comprehensive regression testing. We present Spark-LLM-Eval

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).