DAMamba: Vision State Space Model with Dynamic Adaptive Scan

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tanzhe, Li, Caoshuo, Lyu, Jiayi, Pei, Hongjuan, Zhang, Baochang, Jin, Taisong, Ji, Rongrong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917927436943360
author Li, Tanzhe
Li, Caoshuo
Lyu, Jiayi
Pei, Hongjuan
Zhang, Baochang
Jin, Taisong
Ji, Rongrong
author_facet Li, Tanzhe
Li, Caoshuo
Lyu, Jiayi
Pei, Hongjuan
Zhang, Baochang
Jin, Taisong
Ji, Rongrong
contents State space models (SSMs) have recently garnered significant attention in computer vision. However, due to the unique characteristics of image data, adapting SSMs from natural language processing to computer vision has not outperformed the state-of-the-art convolutional neural networks (CNNs) and Vision Transformers (ViTs). Existing vision SSMs primarily leverage manually designed scans to flatten image patches into sequences locally or globally. This approach disrupts the original semantic spatial adjacency of the image and lacks flexibility, making it difficult to capture complex image structures. To address this limitation, we propose Dynamic Adaptive Scan (DAS), a data-driven method that adaptively allocates scanning orders and regions. This enables more flexible modeling capabilities while maintaining linear computational complexity and global modeling capacity. Based on DAS, we further propose the vision backbone DAMamba, which significantly outperforms current state-of-the-art vision Mamba models in vision tasks such as image classification, object detection, instance segmentation, and semantic segmentation. Notably, it surpasses some of the latest state-of-the-art CNNs and ViTs. Code will be available at https://github.com/ltzovo/DAMamba.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12627
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DAMamba: Vision State Space Model with Dynamic Adaptive Scan
Li, Tanzhe
Li, Caoshuo
Lyu, Jiayi
Pei, Hongjuan
Zhang, Baochang
Jin, Taisong
Ji, Rongrong
Computer Vision and Pattern Recognition
State space models (SSMs) have recently garnered significant attention in computer vision. However, due to the unique characteristics of image data, adapting SSMs from natural language processing to computer vision has not outperformed the state-of-the-art convolutional neural networks (CNNs) and Vision Transformers (ViTs). Existing vision SSMs primarily leverage manually designed scans to flatten image patches into sequences locally or globally. This approach disrupts the original semantic spatial adjacency of the image and lacks flexibility, making it difficult to capture complex image structures. To address this limitation, we propose Dynamic Adaptive Scan (DAS), a data-driven method that adaptively allocates scanning orders and regions. This enables more flexible modeling capabilities while maintaining linear computational complexity and global modeling capacity. Based on DAS, we further propose the vision backbone DAMamba, which significantly outperforms current state-of-the-art vision Mamba models in vision tasks such as image classification, object detection, instance segmentation, and semantic segmentation. Notably, it surpasses some of the latest state-of-the-art CNNs and ViTs. Code will be available at https://github.com/ltzovo/DAMamba.
title DAMamba: Vision State Space Model with Dynamic Adaptive Scan
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.12627