← all papers · overview

Exploring The Reliability Of Large Language Models As Customized Evaluators For Diverse NLP Tasks

Abstract

Previous work adopts large language models (LLMs) as evaluators to evaluate natural language process (NLP) tasks. However, certain shortcomings, e.g., fairness, scope, and accuracy, persist for current LLM evaluators. To analyze whether LLMs can serve as reliable alternatives to humans, we examine the fine-grained alignment between LLM evaluators and human annotators, particularly in understanding

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).