← all papers · overview

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

Abstract

Large reasoning models (LRMs) take longer on harder problems, just as humans do, but that surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong it spends more tokens than when it gets that same problem right; humans do the reverse. We separate two levels of deliberation: how response time tracks difficulty across items (registration), and, with item identity fixed, whether an agent spends more on its own failures or successes (allocation). On a public matched human-LRM corpus, thinking LRMs reproduce the known cross-item alignment with human reaction time but diverge from humans within items: on H-ARC every model lands on the opposite side of zero, the four well-powered ones at Cohen's d = 1.47 to 3.13 against -0.10 for humans. Each agent is scored on its own scale; seconds and tokens never share an axis. The dissociation survives item fixed effects and replicates across datasets; a non-thinking baseline shows a wrong-trial expansion of its own but no cross-item alignment, so part of the effect is not specific to reasoning training. We read the human pattern as engagement versus abandonment: people stay on items they expect to solve and give up on the rest. We read the LRM pattern as length driven by uncertainty: chains grow when the model is unsure, exactly when it tends to fail. Under resource-rational metareasoning these are stopping policies that share a difficulty signal but implement opposite control; trace length captures the signal and misses the control.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).