A Separable Self-attention Inspired by the State Space Model for Computer Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Juntao, Liu, Shaogeng, Bian, Kun, Zhou, You, Zhang, Pei, Liu, Jianning, Zhou, Jun, Liu, Bingyan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915292970483712
author Zhang, Juntao
Liu, Shaogeng
Bian, Kun
Zhou, You
Zhang, Pei
Liu, Jianning
Zhou, Jun
Liu, Bingyan
author_facet Zhang, Juntao
Liu, Shaogeng
Bian, Kun
Zhou, You
Zhang, Pei
Liu, Jianning
Zhou, Jun
Liu, Bingyan
contents Mamba is an efficient State Space Model (SSM) with linear computational complexity. Although SSMs are not suitable for handling non-causal data, Vision Mamba (ViM) methods still demonstrate good performance in tasks such as image classification and object detection. Recent studies have shown that there is a rich theoretical connection between state space models and attention variants. We propose a novel separable self attention method, for the first time introducing some excellent design concepts of Mamba into separable self-attention. To ensure a fair comparison with ViMs, we introduce VMINet, a simple yet powerful prototype architecture, constructed solely by stacking our novel attention modules with the most basic down-sampling layers. Notably, VMINet differs significantly from the conventional Transformer architecture. Our experiments demonstrate that VMINet has achieved competitive results on image classification and high-resolution dense prediction tasks.Code is available at: https://github.com/yws-wxs/VMINet.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02040
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Separable Self-attention Inspired by the State Space Model for Computer Vision
Zhang, Juntao
Liu, Shaogeng
Bian, Kun
Zhou, You
Zhang, Pei
Liu, Jianning
Zhou, Jun
Liu, Bingyan
Computer Vision and Pattern Recognition
Artificial Intelligence
Mamba is an efficient State Space Model (SSM) with linear computational complexity. Although SSMs are not suitable for handling non-causal data, Vision Mamba (ViM) methods still demonstrate good performance in tasks such as image classification and object detection. Recent studies have shown that there is a rich theoretical connection between state space models and attention variants. We propose a novel separable self attention method, for the first time introducing some excellent design concepts of Mamba into separable self-attention. To ensure a fair comparison with ViMs, we introduce VMINet, a simple yet powerful prototype architecture, constructed solely by stacking our novel attention modules with the most basic down-sampling layers. Notably, VMINet differs significantly from the conventional Transformer architecture. Our experiments demonstrate that VMINet has achieved competitive results on image classification and high-resolution dense prediction tasks.Code is available at: https://github.com/yws-wxs/VMINet.
title A Separable Self-attention Inspired by the State Space Model for Computer Vision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2501.02040