Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiong, Zhitong, Wang, Yi, Zhang, Fahong, Stewart, Adam J., Hanna, Joëlle, Borth, Damian, Papoutsis, Ioannis, Saux, Bertrand Le, Camps-Valls, Gustau, Zhu, Xiao Xiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917017183846400
author Xiong, Zhitong
Wang, Yi
Zhang, Fahong
Stewart, Adam J.
Hanna, Joëlle
Borth, Damian
Papoutsis, Ioannis
Saux, Bertrand Le
Camps-Valls, Gustau
Zhu, Xiao Xiang
author_facet Xiong, Zhitong
Wang, Yi
Zhang, Fahong
Stewart, Adam J.
Hanna, Joëlle
Borth, Damian
Papoutsis, Ioannis
Saux, Bertrand Le
Camps-Valls, Gustau
Zhu, Xiao Xiang
contents Earth observation (EO) in open-world settings presents a unique challenge: different applications rely on diverse sensor modalities, each with varying ground sampling distances, spectral ranges, and numbers of spectral bands. However, existing EO foundation models are typically tailored to specific sensor types, making them inflexible when generalizing across the heterogeneous landscape of EO data. To address this, we propose the Dynamic One-For-All (DOFA) model, a unified, multimodal foundation framework designed for diverse vision tasks in EO. Inspired by neural plasticity, DOFA utilizes a wavelength-conditioned dynamic hypernetwork to process inputs from five distinct satellite sensors flexibly. By continually pretraining on five EO modalities, DOFA achieves state-of-the-art performance across multiple downstream tasks and generalizes well to unseen modalities. Enhanced with hybrid continual pretraining, DOFA+ requires significantly fewer computational resources while outperforming counterparts trained with extensive GPU budgets. Experiments on diverse datasets highlight DOFA's potential as a foundation for general-purpose vision models in the sensor-diverse EO domain. The code and pre-trained weights are publicly available at https://github.com/zhu-xlab/DOFA.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15356
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation
Xiong, Zhitong
Wang, Yi
Zhang, Fahong
Stewart, Adam J.
Hanna, Joëlle
Borth, Damian
Papoutsis, Ioannis
Saux, Bertrand Le
Camps-Valls, Gustau
Zhu, Xiao Xiang
Computer Vision and Pattern Recognition
Earth observation (EO) in open-world settings presents a unique challenge: different applications rely on diverse sensor modalities, each with varying ground sampling distances, spectral ranges, and numbers of spectral bands. However, existing EO foundation models are typically tailored to specific sensor types, making them inflexible when generalizing across the heterogeneous landscape of EO data. To address this, we propose the Dynamic One-For-All (DOFA) model, a unified, multimodal foundation framework designed for diverse vision tasks in EO. Inspired by neural plasticity, DOFA utilizes a wavelength-conditioned dynamic hypernetwork to process inputs from five distinct satellite sensors flexibly. By continually pretraining on five EO modalities, DOFA achieves state-of-the-art performance across multiple downstream tasks and generalizes well to unseen modalities. Enhanced with hybrid continual pretraining, DOFA+ requires significantly fewer computational resources while outperforming counterparts trained with extensive GPU budgets. Experiments on diverse datasets highlight DOFA's potential as a foundation for general-purpose vision models in the sensor-diverse EO domain. The code and pre-trained weights are publicly available at https://github.com/zhu-xlab/DOFA.
title Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.15356