RemoteSAM: Towards Segment Anything for Earth Observation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Liang, Liu, Fan, Chen, Delong, Zhang, Chuanyi, Wang, Yijun, Chen, Ziyun, Xu, Wei, Di, Shimin, Zheng, Yuhui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915316960854016
author Yao, Liang
Liu, Fan
Chen, Delong
Zhang, Chuanyi
Wang, Yijun
Chen, Ziyun
Xu, Wei
Di, Shimin
Zheng, Yuhui
author_facet Yao, Liang
Liu, Fan
Chen, Delong
Zhang, Chuanyi
Wang, Yijun
Chen, Ziyun
Xu, Wei
Di, Shimin
Zheng, Yuhui
contents We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github.com/1e12Leon/RemoteSAM.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RemoteSAM: Towards Segment Anything for Earth Observation
Yao, Liang
Liu, Fan
Chen, Delong
Zhang, Chuanyi
Wang, Yijun
Chen, Ziyun
Xu, Wei
Di, Shimin
Zheng, Yuhui
Computer Vision and Pattern Recognition
We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github.com/1e12Leon/RemoteSAM.
title RemoteSAM: Towards Segment Anything for Earth Observation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18022