← all papers · overview

Aqa-bench: An Interactive Benchmark For Evaluating Llms' Sequential Reasoning Ability

Abstract

This paper introduces AQA-Bench, a novel benchmark to assess the sequential reasoning capabilities of large language models (LLMs) in algorithmic contexts, such as depth-first search (DFS). The key feature of our evaluation benchmark lies in its interactive evaluation protocol - for example, in DFS, the availability of each node's connected edge is contingent upon the model's traversal to that nod

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).