← all papers · overview

Reprogramming Vision Foundation Models For Spatio-temporal Forecasting

Abstract

Foundation models have achieved remarkable success in natural language processing and computer vision, demonstrating strong capabilities in modeling complex patterns. While recent efforts have explored adapting large language models (LLMs) for time-series forecasting, LLMs primarily capture one-dimensional sequential dependencies and struggle to model the richer spatio-temporal (ST) correlations essential for accurate ST forecasting. In this paper, we present \textbf\{ST-VFM\}, a novel framework that systematically reprograms Vision Foundation Models (VFMs) for general-purpose spatio-temporal forecasting. While VFMs offer powerful spatial priors, two key challenges arise when applying them to ST tasks: (1) the lack of inherent temporal modeling capacity and (2) the modality gap between visual and ST data. To address these, ST-VFM adopts a *dual-branch architecture* that integrates raw ST inputs with auxiliary ST flow inputs, where the flow encodes lightweight temporal difference signal

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).