Abstract
Bigger language models are less reliable. Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to while within-response knowledge degradation grows up to . We trace that residual to one variable, the per-position disagreement against a stronger oracle, whose second moment splits exactly into bias and decoding risk . That split is an interpretability statement before it is a statistical one: the model's self-readable uncertainty enters only the bias term, so the risk term has no model-readable component. Risk also takes a growing share of the squared error with scale, to from B to B. At a fabrication relaxes within one token while risk persists up to longer, leaving a confident-but-precarious regime that bridges consecutive fabrications ( at B). Contracting that risk at fixed removes - of web-verified hallucinations across six rungs and three families. Semantic entropy fires less on that branch () though it carries nearly the fabrications. Bigger models snowball mistakes faster, through a failure mode that is dominant, self-perpetuating, causal and invisible to the model itself.