RSMamba: Remote Sensing Image Classification with State Space Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Keyan, Chen, Bowen, Liu, Chenyang, Li, Wenyuan, Zou, Zhengxia, Shi, Zhenwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913288893235200
author Chen, Keyan
Chen, Bowen
Liu, Chenyang
Li, Wenyuan
Zou, Zhengxia
Shi, Zhenwei
author_facet Chen, Keyan
Chen, Bowen
Liu, Chenyang
Li, Wenyuan
Zou, Zhengxia
Shi, Zhenwei
contents Remote sensing image classification forms the foundation of various understanding tasks, serving a crucial function in remote sensing image interpretation. The recent advancements of Convolutional Neural Networks (CNNs) and Transformers have markedly enhanced classification accuracy. Nonetheless, remote sensing scene classification remains a significant challenge, especially given the complexity and diversity of remote sensing scenarios and the variability of spatiotemporal resolutions. The capacity for whole-image understanding can provide more precise semantic cues for scene discrimination. In this paper, we introduce RSMamba, a novel architecture for remote sensing image classification. RSMamba is based on the State Space Model (SSM) and incorporates an efficient, hardware-aware design known as the Mamba. It integrates the advantages of both a global receptive field and linear modeling complexity. To overcome the limitation of the vanilla Mamba, which can only model causal sequences and is not adaptable to two-dimensional image data, we propose a dynamic multi-path activation mechanism to augment Mamba's capacity to model non-causal data. Notably, RSMamba maintains the inherent modeling mechanism of the vanilla Mamba, yet exhibits superior performance across multiple remote sensing image classification datasets. This indicates that RSMamba holds significant potential to function as the backbone of future visual foundation models. The code will be available at \url{https://github.com/KyanChen/RSMamba}.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19654
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RSMamba: Remote Sensing Image Classification with State Space Model
Chen, Keyan
Chen, Bowen
Liu, Chenyang
Li, Wenyuan
Zou, Zhengxia
Shi, Zhenwei
Computer Vision and Pattern Recognition
Remote sensing image classification forms the foundation of various understanding tasks, serving a crucial function in remote sensing image interpretation. The recent advancements of Convolutional Neural Networks (CNNs) and Transformers have markedly enhanced classification accuracy. Nonetheless, remote sensing scene classification remains a significant challenge, especially given the complexity and diversity of remote sensing scenarios and the variability of spatiotemporal resolutions. The capacity for whole-image understanding can provide more precise semantic cues for scene discrimination. In this paper, we introduce RSMamba, a novel architecture for remote sensing image classification. RSMamba is based on the State Space Model (SSM) and incorporates an efficient, hardware-aware design known as the Mamba. It integrates the advantages of both a global receptive field and linear modeling complexity. To overcome the limitation of the vanilla Mamba, which can only model causal sequences and is not adaptable to two-dimensional image data, we propose a dynamic multi-path activation mechanism to augment Mamba's capacity to model non-causal data. Notably, RSMamba maintains the inherent modeling mechanism of the vanilla Mamba, yet exhibits superior performance across multiple remote sensing image classification datasets. This indicates that RSMamba holds significant potential to function as the backbone of future visual foundation models. The code will be available at \url{https://github.com/KyanChen/RSMamba}.
title RSMamba: Remote Sensing Image Classification with State Space Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.19654