DiffusionSat: A Generative Foundation Model for Satellite Imagery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khanna, Samar, Liu, Patrick, Zhou, Linqi, Meng, Chenlin, Rombach, Robin, Burke, Marshall, Lobell, David, Ermon, Stefano
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911886572781568
author Khanna, Samar
Liu, Patrick
Zhou, Linqi
Meng, Chenlin
Rombach, Robin
Burke, Marshall
Lobell, David
Ermon, Stefano
author_facet Khanna, Samar
Liu, Patrick
Zhou, Linqi
Meng, Chenlin
Rombach, Robin
Burke, Marshall
Lobell, David
Ermon, Stefano
contents Diffusion models have achieved state-of-the-art results on many modalities including images, speech, and video. However, existing models are not tailored to support remote sensing data, which is widely used in important applications including environmental monitoring and crop-yield prediction. Satellite images are significantly different from natural images -- they can be multi-spectral, irregularly sampled across time -- and existing diffusion models trained on images from the Web do not support them. Furthermore, remote sensing data is inherently spatio-temporal, requiring conditional generation tasks not supported by traditional methods based on captions or images. In this paper, we present DiffusionSat, to date the largest generative foundation model trained on a collection of publicly available large, high-resolution remote sensing datasets. As text-based captions are sparsely available for satellite images, we incorporate the associated metadata such as geolocation as conditioning information. Our method produces realistic samples and can be used to solve multiple generative tasks including temporal generation, superresolution given multi-spectral inputs and in-painting. Our method outperforms previous state-of-the-art methods for satellite image generation and is the first large-scale generative foundation model for satellite imagery. The project website can be found here: https://samar-khanna.github.io/DiffusionSat/
format Preprint
id arxiv_https___arxiv_org_abs_2312_03606
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DiffusionSat: A Generative Foundation Model for Satellite Imagery
Khanna, Samar
Liu, Patrick
Zhou, Linqi
Meng, Chenlin
Rombach, Robin
Burke, Marshall
Lobell, David
Ermon, Stefano
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Diffusion models have achieved state-of-the-art results on many modalities including images, speech, and video. However, existing models are not tailored to support remote sensing data, which is widely used in important applications including environmental monitoring and crop-yield prediction. Satellite images are significantly different from natural images -- they can be multi-spectral, irregularly sampled across time -- and existing diffusion models trained on images from the Web do not support them. Furthermore, remote sensing data is inherently spatio-temporal, requiring conditional generation tasks not supported by traditional methods based on captions or images. In this paper, we present DiffusionSat, to date the largest generative foundation model trained on a collection of publicly available large, high-resolution remote sensing datasets. As text-based captions are sparsely available for satellite images, we incorporate the associated metadata such as geolocation as conditioning information. Our method produces realistic samples and can be used to solve multiple generative tasks including temporal generation, superresolution given multi-spectral inputs and in-painting. Our method outperforms previous state-of-the-art methods for satellite image generation and is the first large-scale generative foundation model for satellite imagery. The project website can be found here: https://samar-khanna.github.io/DiffusionSat/
title DiffusionSat: A Generative Foundation Model for Satellite Imagery
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.03606