Abstract
To evaluate whether LLMs can accurately predict future events, we need the ability to \textit\{backtest\} them on events that have already resolved. This requires models to reason only with information available at a specified past date. Yet LLMs may inadvertently leak post-cutoff knowledge encoded during training, undermining the validity of retrospective evaluation. We introduce a claim-level fr