← all papers · overview

Prism: Dynamic And Flexible Benchmarking Of Llms Code Generation With Monte Carlo Tree Search

Abstract

The rapid advancement of Large Language Models (LLMs) has outpaced traditional evaluation methods. Static benchmarks fail to capture the depth and breadth of LLM capabilities and eventually become obsolete, while most dynamic approaches either rely too heavily on LLM-based evaluation or remain constrained by predefined test sets. We introduce Prism, a flexible, dynamic benchmarking framework desig

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).