← all papers · overview

Lemmabench: A Live, Research-level Benchmark To Evaluate LLM Capabilities In Mathematics

Abstract

We present a new approach for benchmarking Large Language Model (LLM) capabilities on research-level mathematics. Existing benchmarks largely rely on static, hand-curated sets of contest or textbook-style problems as proxies for mathematical research. Instead, we establish an updatable benchmark evaluating models directly on the latest research results in mathematics. This consists of an automatic

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).