Multi-Modal Vision Transformers for Crop Mapping from Satellite Image Time Series

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Follath, Theresa, Mickisch, David, Hemmerling, Jan, Erasmi, Stefan, Schwieder, Marcel, Demir, Begüm
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911930559496192
author Follath, Theresa
Mickisch, David
Hemmerling, Jan
Erasmi, Stefan
Schwieder, Marcel
Demir, Begüm
author_facet Follath, Theresa
Mickisch, David
Hemmerling, Jan
Erasmi, Stefan
Schwieder, Marcel
Demir, Begüm
contents Using images acquired by different satellite sensors has shown to improve classification performance in the framework of crop mapping from satellite image time series (SITS). Existing state-of-the-art architectures use self-attention mechanisms to process the temporal dimension and convolutions for the spatial dimension of SITS. Motivated by the success of purely attention-based architectures in crop mapping from single-modal SITS, we introduce several multi-modal multi-temporal transformer-based architectures. Specifically, we investigate the effectiveness of Early Fusion, Cross Attention Fusion and Synchronized Class Token Fusion within the Temporo-Spatial Vision Transformer (TSViT). Experimental results demonstrate significant improvements over state-of-the-art architectures with both convolutional and self-attention components.
format Preprint
id arxiv_https___arxiv_org_abs_2406_16513
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Modal Vision Transformers for Crop Mapping from Satellite Image Time Series
Follath, Theresa
Mickisch, David
Hemmerling, Jan
Erasmi, Stefan
Schwieder, Marcel
Demir, Begüm
Computer Vision and Pattern Recognition
Using images acquired by different satellite sensors has shown to improve classification performance in the framework of crop mapping from satellite image time series (SITS). Existing state-of-the-art architectures use self-attention mechanisms to process the temporal dimension and convolutions for the spatial dimension of SITS. Motivated by the success of purely attention-based architectures in crop mapping from single-modal SITS, we introduce several multi-modal multi-temporal transformer-based architectures. Specifically, we investigate the effectiveness of Early Fusion, Cross Attention Fusion and Synchronized Class Token Fusion within the Temporo-Spatial Vision Transformer (TSViT). Experimental results demonstrate significant improvements over state-of-the-art architectures with both convolutional and self-attention components.
title Multi-Modal Vision Transformers for Crop Mapping from Satellite Image Time Series
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.16513