Abstract
Large reasoning models (LRMs) take longer on harder problems, just as humans do, but that surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong it spends more tokens than when it gets that same problem right; humans do the reverse. We separate two levels of deliberation: how response time tracks difficulty across items (registration), and, with item identity fixed, whether an agent spends more on its own failures or successes (allocation). On a public matched human-LRM corpus, thinking LRMs reproduce the known cross-item alignment with human reaction time but diverge from humans within items: on H-ARC every model lands on the opposite side of zero, the four well-powered ones at Cohen's d = 1.47 to 3.13 against -0.10 for humans. Each agent is scored on its own scale; seconds and tokens never share an axis. The dissociation survives item fixed effects and replicates across datasets; a non-thinking baseline shows a wrong-trial expansion of its own but no cross-item alignment, so part of the effect is not specific to reasoning training. We read the human pattern as engagement versus abandonment: people stay on items they expect to solve and give up on the rest. We read the LRM pattern as length driven by uncertainty: chains grow when the model is unsure, exactly when it tends to fail. Under resource-rational metareasoning these are stopping policies that share a difficulty signal but implement opposite control; trace length captures the signal and misses the control.