← all papers · overview

Geoeval: Benchmark For Evaluating Llms And Multi-modal Models On Geometry Problem-solving

Abstract

Recent advancements in large language models (LLMs) and multi-modal models (MMs) have demonstrated their remarkable capabilities in problem-solving. Yet, their proficiency in tackling geometry math problems, which necessitates an integrated understanding of both textual and visual information, has not been thoroughly evaluated. To address this gap, we introduce the GeoEval benchmark, a comprehensi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).