← all papers · overview

METAL: Towards Multilingual Meta-evaluation

Abstract

With the rising human-like precision of Large Language Models (LLMs) in numerous tasks, their utilization in a variety of real-world applications is becoming more prevalent. Several studies have shown that LLMs excel on many standard NLP benchmarks. However, it is challenging to evaluate LLMs due to test dataset contamination and the limitations of traditional metrics. Since human evaluations are

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).