Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Perera, Amal S., Fernandez, David, Witharana, Chandi, Manos, Elias, Pimenta, Michael, Liljedahl, Anna K., Nitze, Ingmar, Yang, Yili, Nicholson, Todd, Hsu, Chia-Yu, Li, Wenwen, Grosse, Guido
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915321450856448
author Perera, Amal S.
Fernandez, David
Witharana, Chandi
Manos, Elias
Pimenta, Michael
Liljedahl, Anna K.
Nitze, Ingmar
Yang, Yili
Nicholson, Todd
Hsu, Chia-Yu
Li, Wenwen
Grosse, Guido
author_facet Perera, Amal S.
Fernandez, David
Witharana, Chandi
Manos, Elias
Pimenta, Michael
Liljedahl, Anna K.
Nitze, Ingmar
Yang, Yili
Nicholson, Todd
Hsu, Chia-Yu
Li, Wenwen
Grosse, Guido
contents Accurate mapping of permafrost landforms, thaw disturbances, and human-built infrastructure at pan-Arctic scale using sub-meter satellite imagery is increasingly critical. Handling petabyte-scale image data requires high-performance computing and robust feature detection models. While convolutional neural network (CNN)-based deep learning approaches are widely used for remote sensing (RS),similar to the success in transformer based large language models, Vision Transformers (ViTs) offer advantages in capturing long-range dependencies and global context via attention mechanisms. ViTs support pretraining via self-supervised learning-addressing the common limitation of labeled data in Arctic feature detection and outperform CNNs on benchmark datasets. Arctic also poses challenges for model generalization, especially when features with the same semantic class exhibit diverse spectral characteristics. To address these issues for Arctic feature detection, we integrate geospatial location embeddings into ViTs to improve adaptation across regions. This work investigates: (1) the suitability of pre-trained ViTs as feature extractors for high-resolution Arctic remote sensing tasks, and (2) the benefit of combining image and location embeddings. Using previously published datasets for Arctic feature detection, we evaluate our models on three tasks-detecting ice-wedge polygons (IWP), retrogressive thaw slumps (RTS), and human-built infrastructure. We empirically explore multiple configurations to fuse image embeddings and location embeddings. Results show that ViTs with location embeddings outperform prior CNN-based models on two of the three tasks including F1 score increase from 0.84 to 0.92 for RTS detection, demonstrating the potential of transformer-based models with spatial awareness for Arctic RS applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02868
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
Perera, Amal S.
Fernandez, David
Witharana, Chandi
Manos, Elias
Pimenta, Michael
Liljedahl, Anna K.
Nitze, Ingmar
Yang, Yili
Nicholson, Todd
Hsu, Chia-Yu
Li, Wenwen
Grosse, Guido
Computer Vision and Pattern Recognition
I.4.6; I.5.4; I.5.2; I.2.10
Accurate mapping of permafrost landforms, thaw disturbances, and human-built infrastructure at pan-Arctic scale using sub-meter satellite imagery is increasingly critical. Handling petabyte-scale image data requires high-performance computing and robust feature detection models. While convolutional neural network (CNN)-based deep learning approaches are widely used for remote sensing (RS),similar to the success in transformer based large language models, Vision Transformers (ViTs) offer advantages in capturing long-range dependencies and global context via attention mechanisms. ViTs support pretraining via self-supervised learning-addressing the common limitation of labeled data in Arctic feature detection and outperform CNNs on benchmark datasets. Arctic also poses challenges for model generalization, especially when features with the same semantic class exhibit diverse spectral characteristics. To address these issues for Arctic feature detection, we integrate geospatial location embeddings into ViTs to improve adaptation across regions. This work investigates: (1) the suitability of pre-trained ViTs as feature extractors for high-resolution Arctic remote sensing tasks, and (2) the benefit of combining image and location embeddings. Using previously published datasets for Arctic feature detection, we evaluate our models on three tasks-detecting ice-wedge polygons (IWP), retrogressive thaw slumps (RTS), and human-built infrastructure. We empirically explore multiple configurations to fuse image embeddings and location embeddings. Results show that ViTs with location embeddings outperform prior CNN-based models on two of the three tasks including F1 score increase from 0.84 to 0.92 for RTS detection, demonstrating the potential of transformer-based models with spatial awareness for Arctic RS applications.
title Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
topic Computer Vision and Pattern Recognition
I.4.6; I.5.4; I.5.2; I.2.10
url https://arxiv.org/abs/2506.02868