Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing Images

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zou, Xuechao, Zhang, Shun, Li, Kai, Wang, Shiying, Xing, Junliang, Jin, Lei, Lang, Congyan, Tao, Pin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910712123621376
author Zou, Xuechao
Zhang, Shun
Li, Kai
Wang, Shiying
Xing, Junliang
Jin, Lei
Lang, Congyan
Tao, Pin
author_facet Zou, Xuechao
Zhang, Shun
Li, Kai
Wang, Shiying
Xing, Junliang
Jin, Lei
Lang, Congyan
Tao, Pin
contents Cloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated powerful generalization capabilities across various visual tasks. In this paper, we present a parameter-efficient adaptive approach, termed Cloud-Adapter, designed to enhance the accuracy and robustness of cloud segmentation. Our method leverages a VFM pretrained on general domain data, which remains frozen, eliminating the need for additional training. Cloud-Adapter incorporates a lightweight spatial perception module that initially utilizes a convolutional neural network (ConvNet) to extract dense spatial representations. These multi-scale features are then aggregated and serve as contextual inputs to an adapting module, which modulates the frozen transformer layers within the VFM. Experimental results demonstrate that the Cloud-Adapter approach, utilizing only 0.6% of the trainable parameters of the frozen backbone, achieves substantial performance gains. Cloud-Adapter consistently achieves state-of-the-art performance across various cloud segmentation datasets from multiple satellite sources, sensor series, data processing levels, land cover scenarios, and annotation granularities. We have released the code and model checkpoints at https://xavierjiezou.github.io/Cloud-Adapter/ to support further research.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13127
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing Images
Zou, Xuechao
Zhang, Shun
Li, Kai
Wang, Shiying
Xing, Junliang
Jin, Lei
Lang, Congyan
Tao, Pin
Computer Vision and Pattern Recognition
Cloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated powerful generalization capabilities across various visual tasks. In this paper, we present a parameter-efficient adaptive approach, termed Cloud-Adapter, designed to enhance the accuracy and robustness of cloud segmentation. Our method leverages a VFM pretrained on general domain data, which remains frozen, eliminating the need for additional training. Cloud-Adapter incorporates a lightweight spatial perception module that initially utilizes a convolutional neural network (ConvNet) to extract dense spatial representations. These multi-scale features are then aggregated and serve as contextual inputs to an adapting module, which modulates the frozen transformer layers within the VFM. Experimental results demonstrate that the Cloud-Adapter approach, utilizing only 0.6% of the trainable parameters of the frozen backbone, achieves substantial performance gains. Cloud-Adapter consistently achieves state-of-the-art performance across various cloud segmentation datasets from multiple satellite sources, sensor series, data processing levels, land cover scenarios, and annotation granularities. We have released the code and model checkpoints at https://xavierjiezou.github.io/Cloud-Adapter/ to support further research.
title Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.13127