← all papers · overview

Beyond Accuracy: Evaluating The Reasoning Behavior Of Large Language Models -- A Survey

Abstract

Large language models (LLMs) have recently shown impressive performance on tasks involving reasoning, leading to a lively debate on whether these models possess reasoning capabilities similar to humans. However, despite these successes, the depth of LLMs' reasoning abilities remains uncertain. This uncertainty partly stems from the predominant focus on task performance, measured through shallow ac

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).