← all papers · overview

Docpuzzle: A Process-aware Benchmark For Evaluating Realistic Long-context Reasoning Capabilities

Abstract

We present DocPuzzle, a rigorously constructed benchmark for evaluating long-context reasoning capabilities in large language models (LLMs). This benchmark comprises 100 expert-level QA problems requiring multi-step reasoning over long real-world documents. To ensure the task quality and complexity, we implement a human-AI collaborative annotation-validation pipeline. DocPuzzle introduces an innov

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).