Abstract
We investigate whether structured reasoning interventions improve the strategic economic reasoning of large language models, and whether their effects depend on model architecture. Using Hotelling's linear city model as a diagnostic vehicle, we evaluate GPT-4.1-mini (a standard instruction-following model) and GPT-5-mini (a reasoning-optimized model) under five conditions - an unscaffolded baseline and four reasoning interventions - across eight questions spanning deductive and abductive reasoning, three prompt framings, and three repetitions per condition, yielding 720 individually judged responses. We find a statistically significant crossover interaction between scaffolding type and model architecture (, , ): commitment scaffolding improves the standard model () while degrading the reasoning model (), and principled separation shows the opposite pattern ( vs. ). Both crossovers are individually significant (commitment: ; separation: ) and hold across all eight questions with 7/8 directional consistency. Adversarial stress-testing harms both models, with greater degradation for the reasoning model ( vs. ; ), and the damage correlates negatively with baseline difficulty (, ). We further document a persistent declarative-procedural gap in which both models identify correct strategies at rates far exceeding their ability to execute them; separation fully closes this gap for the reasoning model while no intervention helps the standard model.