SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jain, Pallavi, Ienco, Dino, Interdonato, Roberto, Berchoux, Tristan, Marcos, Diego
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912538420051968
author Jain, Pallavi
Ienco, Dino
Interdonato, Roberto
Berchoux, Tristan
Marcos, Diego
author_facet Jain, Pallavi
Ienco, Dino
Interdonato, Roberto
Berchoux, Tristan
Marcos, Diego
contents Pre-trained vision-language models (VLMs), such as CLIP, demonstrate impressive zero-shot classification capabilities with free-form prompts and even show some generalization in specialized domains. However, their performance on satellite imagery is limited due to the underrepresentation of such data in their training sets, which predominantly consist of ground-level images. Existing prompting techniques for satellite imagery are often restricted to generic phrases like a satellite image of ..., limiting their effectiveness for zero-shot land-use and land-cover (LULC) mapping. To address these challenges, we introduce SenCLIP, which transfers CLIPs representation to Sentinel-2 imagery by leveraging a large dataset of Sentinel-2 images paired with geotagged ground-level photos from across Europe. We evaluate SenCLIP alongside other SOTA remote sensing VLMs on zero-shot LULC mapping tasks using the EuroSAT and BigEarthNet datasets with both aerial and ground-level prompting styles. Our approach, which aligns ground-level representations with satellite imagery, demonstrates significant improvements in classification accuracy across both prompt styles, opening new possibilities for applying free-form textual descriptions in zero-shot LULC mapping.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08536
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting
Jain, Pallavi
Ienco, Dino
Interdonato, Roberto
Berchoux, Tristan
Marcos, Diego
Computer Vision and Pattern Recognition
Pre-trained vision-language models (VLMs), such as CLIP, demonstrate impressive zero-shot classification capabilities with free-form prompts and even show some generalization in specialized domains. However, their performance on satellite imagery is limited due to the underrepresentation of such data in their training sets, which predominantly consist of ground-level images. Existing prompting techniques for satellite imagery are often restricted to generic phrases like a satellite image of ..., limiting their effectiveness for zero-shot land-use and land-cover (LULC) mapping. To address these challenges, we introduce SenCLIP, which transfers CLIPs representation to Sentinel-2 imagery by leveraging a large dataset of Sentinel-2 images paired with geotagged ground-level photos from across Europe. We evaluate SenCLIP alongside other SOTA remote sensing VLMs on zero-shot LULC mapping tasks using the EuroSAT and BigEarthNet datasets with both aerial and ground-level prompting styles. Our approach, which aligns ground-level representations with satellite imagery, demonstrates significant improvements in classification accuracy across both prompt styles, opening new possibilities for applying free-form textual descriptions in zero-shot LULC mapping.
title SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.08536