A Recipe For Improving Remote Sensing VLM Zero Shot Generalization
2025 Β· Aviad Barzilai, Yotam Gigi, Amr Helmy, et al.
Abstract
Foundation models have had a significant impact across various AI applications, enabling use cases that were previously impossible. Contrastive Visual Language Models (VLMs), in particular, have outperformed other techniques in many tasks. However, their prevalence in remote sensing (RS) is still limited, due to the scarcity of diverse remote-sensing visual-language datasets. In this work we introduce two novel image-caption datasets for training of remote sensing foundation models. The first dataset pairs aerial and satellite imagery with captions generated by Gemini using landmarks extracted from Google Maps. The second dataset utilizes public web images and their corresponding alt-text, filtered for the remote sensing domain, resulting in a diverse dataset with greater breadth in image styles and subject matter. These datasets are used to pre-train the MaMMUT~\citep\{kuo2023mammutsimplearchitecturejoint\} VLM architecture, resulting in state-of-the-art generalization performance in
Authors
(none)
Tags
Stats
Related papers
- Remote Sensing Retrieval-augmented Generation: Bridging Remote Sensing Imagery And Comprehensive Knowledge With A Multi-modal Dataset And Retrieval-augmented Generation Model (2025)2.26
- Redundancy-aware Pretraining Of Vision-language Foundation Models In Remote Sensing (2025)2.26
- Vlm2geovec: Toward Universal Multimodal Embeddings For Remote Sensing (2025)0.00
- Large Language Models For Captioning And Retrieving Remote Sensing Images (2024)0.00
- DGTRSD & DGTRS-CLIP: A Dual-granularity Remote Sensing Image-text Dataset And Vision Language Foundation Model For Alignment (2025)2.98
- Vision-language Modelling For Radiological Imaging And Reports In The Low Data Regime (2023)0.00
- Mlrsnet: A Multi-label High Spatial Resolution Remote Sensing Dataset For Semantic Scene Understanding (2020)19.63
- Multi-spectral Remote Sensing Image Retrieval Using Geospatial Foundation Models (2024)9.27