← all papers · overview

Evaluating Large Language Models With Grid-based Game Competitions: An Extensible LLM Benchmark And Leaderboard

Abstract

We introduce a novel and extensible benchmark for large language models (LLMs) through grid-based games such as Tic-Tac-Toe, Connect Four, and Gomoku. The open-source game simulation code, available on GitHub, allows LLMs to compete and generates detailed data files in JSON, CSV, TXT, and PNG formats for leaderboard rankings and further analysis. We present the results of games among leading LLMs,

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).