← all papers · overview

Cer-eval: Certifiable And Cost-efficient Evaluation Framework For Llms

Abstract

As foundation models continue to scale, the size of trained models grows exponentially, presenting significant challenges for their evaluation. Current evaluation practices involve curating increasingly large datasets to assess the performance of large language models (LLMs). However, there is a lack of systematic analysis and guidance on determining the sufficiency of test data or selecting infor

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).