← all papers · overview

Botzonebench: Scalable LLM Evaluation Via Graded AI Anchors

Abstract

Large Language Models (LLMs) are increasingly deployed in interactive environments requiring strategic decision-making, yet systematic evaluation of these capabilities remains challenging. Existing benchmarks for LLMs primarily assess static reasoning through isolated tasks and fail to capture dynamic strategic abilities. Recent game-based evaluations employ LLM-vs-LLM tournaments that produce rel

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).