Abstract
Reference-free evaluation of large language model (LLM) creativity relies on perplexity, entropy, and top-1 margin. We show that a much stronger signal lives one step earlier in the pipeline: in how sampling temperature \emph{reshapes} the model's token distribution before the next token is drawn. On Llama-3.1-8B-Instruct generations of 500 open-ended creative prompts at , a single per-token feature derived from this reshaping predicts the within-prompt creativity rank at Spearman against an averaged gpt-4o\,/\,gemini-2.5-pro judge () and against a three-rater human-majority ranking (). Each of four standard reference-free baselines (self-perplexity, mean predictive entropy, top-1 margin, gzip compression ratio) tops out at on both ground truths: a gap of on averaged-LLM and on human-majority, both far larger than the spread among the baselines themselves. The two ground-truth panels agree with each other at , above the inter-human ceiling of , so the comparison is not bottlenecked by judge noise. Mechanistically, the win comes from a sharp distributional signature of the incoherence regime: at the cumulative-mass width inflates from to tokens and post-temperature mass leaks off the pre-temperature top- plausible set by about percentage points. The per-token aggregates do not separate from ; discriminating the two coherent regimes is left to sequence-level features.