DefMamba: Deformable Visual State Space Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Leiye, Zhang, Miao, Yin, Jihao, Liu, Tingwei, Ji, Wei, Piao, Yongri, Lu, Huchuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916678746505216
author Liu, Leiye
Zhang, Miao
Yin, Jihao
Liu, Tingwei
Ji, Wei
Piao, Yongri
Lu, Huchuan
author_facet Liu, Leiye
Zhang, Miao
Yin, Jihao
Liu, Tingwei
Ji, Wei
Piao, Yongri
Lu, Huchuan
contents Recently, state space models (SSM), particularly Mamba, have attracted significant attention from scholars due to their ability to effectively balance computational efficiency and performance. However, most existing visual Mamba methods flatten images into 1D sequences using predefined scan orders, which results the model being less capable of utilizing the spatial structural information of the image during the feature extraction process. To address this issue, we proposed a novel visual foundation model called DefMamba. This model includes a multi-scale backbone structure and deformable mamba (DM) blocks, which dynamically adjust the scanning path to prioritize important information, thus enhancing the capture and processing of relevant input features. By combining a deformable scanning(DS) strategy, this model significantly improves its ability to learn image structures and detects changes in object details. Numerous experiments have shown that DefMamba achieves state-of-the-art performance in various visual tasks, including image classification, object detection, instance segmentation, and semantic segmentation. The code is open source on DefMamba.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DefMamba: Deformable Visual State Space Model
Liu, Leiye
Zhang, Miao
Yin, Jihao
Liu, Tingwei
Ji, Wei
Piao, Yongri
Lu, Huchuan
Computer Vision and Pattern Recognition
Recently, state space models (SSM), particularly Mamba, have attracted significant attention from scholars due to their ability to effectively balance computational efficiency and performance. However, most existing visual Mamba methods flatten images into 1D sequences using predefined scan orders, which results the model being less capable of utilizing the spatial structural information of the image during the feature extraction process. To address this issue, we proposed a novel visual foundation model called DefMamba. This model includes a multi-scale backbone structure and deformable mamba (DM) blocks, which dynamically adjust the scanning path to prioritize important information, thus enhancing the capture and processing of relevant input features. By combining a deformable scanning(DS) strategy, this model significantly improves its ability to learn image structures and detects changes in object details. Numerous experiments have shown that DefMamba achieves state-of-the-art performance in various visual tasks, including image classification, object detection, instance segmentation, and semantic segmentation. The code is open source on DefMamba.
title DefMamba: Deformable Visual State Space Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.05794