Abstract
Reasoning with a chain-of-thought (CoT) enables Large Language Models (LLMs) to solve complex tasks but incurs significant inference costs due to the generation of long rationales. We propose Thinking States, a method that performs reasoning \{\em while\} the input is processing. Specifically, Thinking States generates sequences of thinking tokens every few input tokens, transforms the thoughts ba