← all papers · overview

EQUATOR: A Deterministic Framework For Evaluating LLM Reasoning With Open-ended Questions. # V1.0.0-beta

Abstract

Despite the remarkable coherence of Large Language Models (LLMs), existing evaluation methods often suffer from fluency bias and rely heavily on multiple-choice formats, making it difficult to assess factual accuracy and complex reasoning effectively. LLMs thus frequently generate factually inaccurate responses, especially in complex reasoning tasks, highlighting two prominent challenges: (1) the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).