SF-Mamba: Rethinking State Space Model for Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yoshimura, Masakazu, Hayashi, Teruaki, Hoshino, Yuki, Wang, Wei-Yao, Ohashi, Takeshi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908893790076928
author Yoshimura, Masakazu
Hayashi, Teruaki
Hoshino, Yuki
Wang, Wei-Yao
Ohashi, Takeshi
author_facet Yoshimura, Masakazu
Hayashi, Teruaki
Hoshino, Yuki
Wang, Wei-Yao
Ohashi, Takeshi
contents The realm of Mamba for vision has been advanced in recent years to strike for the alternatives of Vision Transformers (ViTs) that suffer from the quadratic complexity. While the recurrent scanning mechanism of Mamba offers computational efficiency, it inherently limits non-causal interactions between image patches. Prior works have attempted to address this limitation through various multi-scan strategies; however, these approaches suffer from inefficiencies due to suboptimal scan designs and frequent data rearrangement. Moreover, Mamba exhibits relatively slow computational speed under short token lengths, commonly used in visual tasks. In pursuit of a truly efficient vision encoder, we rethink the scan operation for vision and the computational efficiency of Mamba. To this end, we propose SF-Mamba, a novel visual Mamba with two key proposals: auxiliary patch swapping for encoding bidirectional information flow under an unidirectional scan and batch folding with periodic state reset for advanced GPU parallelism. Extensive experiments on image classification, object detection, and instance and semantic segmentation consistently demonstrate that our proposed SF-Mamba significantly outperforms state-of-the-art baselines while improving throughput across different model sizes. We will release the source code after publication.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16423
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SF-Mamba: Rethinking State Space Model for Vision
Yoshimura, Masakazu
Hayashi, Teruaki
Hoshino, Yuki
Wang, Wei-Yao
Ohashi, Takeshi
Computer Vision and Pattern Recognition
Artificial Intelligence
The realm of Mamba for vision has been advanced in recent years to strike for the alternatives of Vision Transformers (ViTs) that suffer from the quadratic complexity. While the recurrent scanning mechanism of Mamba offers computational efficiency, it inherently limits non-causal interactions between image patches. Prior works have attempted to address this limitation through various multi-scan strategies; however, these approaches suffer from inefficiencies due to suboptimal scan designs and frequent data rearrangement. Moreover, Mamba exhibits relatively slow computational speed under short token lengths, commonly used in visual tasks. In pursuit of a truly efficient vision encoder, we rethink the scan operation for vision and the computational efficiency of Mamba. To this end, we propose SF-Mamba, a novel visual Mamba with two key proposals: auxiliary patch swapping for encoding bidirectional information flow under an unidirectional scan and batch folding with periodic state reset for advanced GPU parallelism. Extensive experiments on image classification, object detection, and instance and semantic segmentation consistently demonstrate that our proposed SF-Mamba significantly outperforms state-of-the-art baselines while improving throughput across different model sizes. We will release the source code after publication.
title SF-Mamba: Rethinking State Space Model for Vision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.16423