← all papers · overview

Who Benchmarks The Benchmarks? A Case Study Of LLM Evaluation In Icelandic

Abstract

This paper evaluates current Large Language Model (LLM) benchmarking for Icelandic, identifies problems, and calls for improved evaluation methods in low/medium-resource languages in particular. We show that benchmarks that include synthetic or machine-translated data that have not been verified in any way, commonly contain severely flawed test examples that are likely to skew the results and unde

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).