← all papers · overview

Competition-level Problems Are Effective LLM Evaluators

Abstract

Large language models (LLMs) have demonstrated impressive reasoning capabilities, yet there is ongoing debate about these abilities and the potential data contamination problem recently. This paper aims to evaluate the reasoning capacities of LLMs, specifically in solving recent competition-level programming problems in Codeforces, which are expert-crafted and unique, requiring deep understanding

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).