← all papers · overview

PARIKSHA: A Large-scale Investigation Of Human-llm Evaluator Agreement On Multilingual And Multi-cultural Data

Abstract

Evaluation of multilingual Large Language Models (LLMs) is challenging due to a variety of factors -- the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and the lack of local, cultural nuances in translated benchmarks. In this work, we study human and LLM-based evaluation in a multilingual, multi-cultural setting. We evaluate

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).