Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 11:00 PM ET on Thursday, August 13 until 12:00 AM ET on Friday, August 14 due to maintenance. We apologize for the inconvenience.


Title: Self-Supervised Learning across the Spectrum
Satellite image time series (SITS) segmentation is crucial for many applications, like environmental monitoring, land cover mapping, and agricultural crop type classification. However, training models for SITS segmentation remains a challenging task due to the lack of abundant training data, which requires fine-grained annotation. We propose S4, a new self-supervised pretraining approach that significantly reduces the requirement for labeled training data by utilizing two key insights of satellite imagery: (a) Satellites capture images in different parts of the spectrum, such as radio frequencies and visible frequencies. (b) Satellite imagery is geo-registered, allowing for fine-grained spatial alignment. We use these insights to formulate pretraining tasks in S4. To the best of our knowledge, S4 is the first multimodal and temporal approach for SITS segmentation. S4’s novelty stems from leveraging multiple properties required for SITS self-supervision: (1) multiple modalities, (2) temporal information, and (3) pixel-level feature extraction. We also curate m2s2-SITS, a large-scale dataset of unlabeled, spatially aligned, multimodal, and geographic-specific SITS that serves as representative pretraining data for S4. Finally, we evaluate S4 on multiple SITS segmentation datasets and demonstrate its efficacy against competing baselines while using limited labeled data. Through a series of extensive comparisons and ablation studies, we demonstrate S4’s ability as an effective feature extractor for downstream semantic segmentation.  more » « less
Award ID(s):
2237474
PAR ID:
10577830
Author(s) / Creator(s):
; ; ; ; ; ;
Publisher / Repository:
MDPI Remote Sensing
Date Published:
Journal Name:
Remote Sensing
Volume:
16
Issue:
18
ISSN:
2072-4292
Page Range / eLocation ID:
3470
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract The world’s coastlines are spatially highly variable, coupled-human-natural systems that comprise a nested hierarchy of component landforms, ecosystems, and human interventions, each interacting over a range of space and time scales. Understanding and predicting coastline dynamics necessitates frequent observation from imaging sensors on remote sensing platforms. Machine Learning models that carry out supervised (i.e., human-guided) pixel-based classification, or image segmentation, have transformative applications in spatio-temporal mapping of dynamic environments, including transient coastal landforms, sediments, habitats, waterbodies, and water flows. However, these models require large and well-documented training and testing datasets consisting of labeled imagery. We describe “Coast Train,” a multi-labeler dataset of orthomosaic and satellite images of coastal environments and corresponding labels. These data include imagery that are diverse in space and time, and contain 1.2 billion labeled pixels, representing over 3.6 million hectares. We use a human-in-the-loop tool especially designed for rapid and reproducible Earth surface image segmentation. Our approach permits image labeling by multiple labelers, in turn enabling quantification of pixel-level agreement over individual and collections of images. 
    more » « less
  2. Despite recent progress in computer vision, fine-grained interpretation of satellite images remains challenging because of a lack of labeled training data. To overcome this limitation, we construct a novel dataset called WikiSatNet by pairing geo-referenced Wikipedia articles with satellite imagery of their corresponding locations. We then propose two strategies to learn representations of satellite images by predicting properties of the corresponding articles from the images. Leveraging this new multi-modal dataset, we can drastically reduce the quantity of human-annotated labels and time required for downstream tasks. On the recently released fMoW dataset, our pre-training strategies can boost the performance of a model pre-trained on ImageNet by up to 4.5% in F1 score. 
    more » « less
  3. null (Ed.)
    Fine-scale sea ice conditions are key to our efforts to understand and model climate change. We propose the first deep learning pipeline to extract fine-scale sea ice layers from high-resolution satellite imagery (Worldview-3). Extracting sea ice from imagery is often challenging due to the potentially complex texture from older ice floes (i.e., floating chunks of sea ice) and surrounding slush ice, making ice floes less distinctive from the surrounding water. We propose a pipeline using a U-Net variant with a Resnet encoder to retrieve ice floe pixel masks from very-high-resolution multispectral satellite imagery. Even with a modest-sized hand-labeled training set and the most basic hyperparameter choices, our CNN-based approach attains an out-of-sample F1 score of 0.698–a nearly 60% improvement when compared to a watershed segmentation baseline. We then supplement our training set with a much larger sample of images weak-labeled by a watershed segmentation algorithm. To ensure watershed derived pack-ice masks were a good representation of the underlying images, we created a synthetic version for each weak-labeled image, where areas outside the mask are replaced by open water scenery. Adding our synthetic image dataset, obtained at minimal effort when compared with hand-labeling, further improves the out-of-sample F1 score to 0.734. Finally, we use an ensemble of four test metrics and evaluated after mosaicing outputs for entire scenes to mimic production setting during model selection, reaching an out-of-sample F1 score of 0.753. Our fully-automated pipeline is capable of detecting, monitoring, and segmenting ice floes at a very fine level of detail, and provides a roadmap for other use-cases where partial results can be obtained with threshold-based methods but a context-robust segmentation pipeline is desired. 
    more » « less
  4. To date, Deep Learning models for archaeological feature detection have generally been built on the back of off-the-shelf convolutional neural networks (CNNs) and vision Transformer (ViT) models, which are pretrained on a variety of image types, sources, and subjects that are not specific to analyzing high-resolution satellite imagery. Recent advances in transformer-based vision models and self-supervised training approaches make it possible for researchers to generate foundation models that are more finely attuned to specific domains, without huge amounts of human-annotated training data. We discuss the development of two such models employing Meta's transformer-based DINOv2 framework. The first, DeepAndes, is based on the ingestion of a 3 million chip sample from a two million square km area of high-resolution multispectral satellite imagery of the Andean region. This foundation model has broad utility across the social and earth sciences. The second, DeepAndesArch is fine-tuned labeled archaeological training data collected by the GeoPACHA project to create an archaeology-focused version of DeepAndes. We present the processes involved in generating DeepAndes and DeepAndesArch and discuss prospects for foundation models in archaeological research 
    more » « less
  5. Research employing AI and Deep Learning for archaeological survey of satellite imagery has expanded rapidly over the past five years (e.g., Fraser et al. 2024; Kadhim and Abed 2023; Landauer and Klassen 2025; Orengo et al. 2020). However, the scale of archaeological data tends to remain relatively small by the standard of AI applications. ImageNet––perhaps the best known benchmark dataset in Computer Vision research––contains more than 14 million images. In contrast, even the most intensive satellite imagery survey projects in archaeology have tended to identify loci numbering, at most, in the 10s of thousands of features (Wernke et al. 2024)––comparable to the number of images in the small-scale computer vision dataset MNIST, which is used in introductory machine learning classes. Indeed, more typically, archaeological satellite survey data operates on the scale of dozens or hundreds of features, such as a recent publication analyzing the regional distribution of 76 Andean hunting traps (Oyaneder 2025). Of course, due to cultural and chronological variation, modeling archaeological sites in satellite data is much more challenging than distinguishing between, say, cats and dogs––or between standardized models of automobiles. As a result, we do not have sufficient data for effectively training traditional deep learning models. A growing number of studies have shown that training AI models to recognize specific archaeological features requires large, feature-specific labeled datasets. For example, Castiello et al (2025)'s note that more than one thousand annotated examples of circular structures ("pucaras") were required in order to train a CNN capable of reliably identifying them in the desert landscape of northern Chile. However, archaeologists are often interested in many different feature types and scales (including highly idiosyncratic feature types), which makes it impractical to create separate large annotated datasets and models for each individual feature of interest. Moreover, the lack of significant profit incentives for heritage applications means that there have been few large-scale attempts to create training data for archaeological feature detection. Recent generalizable vision foundation models (VFMs) such as DINOv3, AlphaEarth, and DeepAndes can help us address these issues by enabling few shot and zero shot feature detection and segmentation. These VFMs are pretrained on extensive datasets. For example, DINOv3 is trained on over 1 billion images, while our Andes-specific model DeepAndes is pretrained on more than 3 million multispectral WorldView satellite images (Guo et al. 2025). Through large-scale self-supervised learning, these models learn highly transferable visual representations that can be adapted to new feature types (e.g., archaeological loci) with minimal additional annotation. As a result, they enable few-shot and even zero-shot computer vision tasks in archaeological remote sensing, reducing the need to build a dedicated large labeled dataset for every feature of archaeological interest. For example, archaeologists can provide a small but diverse set of example images of a particular archaeological feature. The VFM embeds these examples as anchor representations in its feature space and can then retrieve visually similar loci across large regions without additional training (i.e., in a zero-shot manner). This allows researchers to rapidly identify new and previously undocumented instances of a feature at continental scale—an otherwise infeasible task using manual methods—and to better understand its spatial distribution and associated cultural or environmental patterns. For few-shot detection and segmentation, the adaptable backbones (usually Vision Transformers) of VFMs can be plugged into existing, well-established computer vision frameworks. For example, by integrating these backbones into state-of-the-art detection models such as YOLO, we can enable low-resource fine-tuning and provide archaeologists with practical tools for detecting archaeological loci using bounding boxes. This makes it possible to develop effective detection systems even when only a small number of labeled examples are available. In turn, this accelerates the pace of archaeological survey, helping to reveal broader patterns of settlement, infrastructure, and cultural interaction. Through these means, we leverage our improved VFM DeepAndes 2.0, and, together with general-purpose vision foundation models such as DINOv3, we accelerate and scale our archaeological survey of the central Andes. The examples retrieved through large-scale similarity search are then used to train feature-specific detection models, enabling the efficient and scalable identification of archaeological features such as corrals, buildings, and terracings. 
    more » « less