A Survey on Mamba Architecture for Vision Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ibrahim, Fady, Liu, Guangjun, Wang, Guanghui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909487989784576
author Ibrahim, Fady
Liu, Guangjun
Wang, Guanghui
author_facet Ibrahim, Fady
Liu, Guangjun
Wang, Guanghui
contents Transformers have become foundational for visual tasks such as object detection, semantic segmentation, and video understanding, but their quadratic complexity in attention mechanisms presents scalability challenges. To address these limitations, the Mamba architecture utilizes state-space models (SSMs) for linear scalability, efficient processing, and improved contextual awareness. This paper investigates Mamba architecture for visual domain applications and its recent advancements, including Vision Mamba (ViM) and VideoMamba, which introduce bidirectional scanning, selective scanning mechanisms, and spatiotemporal processing to enhance image and video understanding. Architectural innovations like position embeddings, cross-scan modules, and hierarchical designs further optimize the Mamba framework for global and local feature extraction. These advancements position Mamba as a promising architecture in computer vision research and applications.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07161
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey on Mamba Architecture for Vision Applications
Ibrahim, Fady
Liu, Guangjun
Wang, Guanghui
Computer Vision and Pattern Recognition
Artificial Intelligence
Transformers have become foundational for visual tasks such as object detection, semantic segmentation, and video understanding, but their quadratic complexity in attention mechanisms presents scalability challenges. To address these limitations, the Mamba architecture utilizes state-space models (SSMs) for linear scalability, efficient processing, and improved contextual awareness. This paper investigates Mamba architecture for visual domain applications and its recent advancements, including Vision Mamba (ViM) and VideoMamba, which introduce bidirectional scanning, selective scanning mechanisms, and spatiotemporal processing to enhance image and video understanding. Architectural innovations like position embeddings, cross-scan modules, and hierarchical designs further optimize the Mamba framework for global and local feature extraction. These advancements position Mamba as a promising architecture in computer vision research and applications.
title A Survey on Mamba Architecture for Vision Applications
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2502.07161