← all papers · overview

Gameeval: Evaluating Llms On Conversational Games

Abstract

The rapid advancements in large language models (LLMs) have presented challenges in evaluating those models. Existing evaluation methods are either reference-based or preference based, which inevitably need human intervention or introduce test bias caused by evaluator models. In this paper, we propose GameEval, a novel approach to evaluating LLMs through goal-driven conversational games, overcomin

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).