Mamba in Vision: A Comprehensive Survey of Techniques and Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rahman, Md Maklachur, Tutul, Abdullah Aman, Nath, Ankur, Laishram, Lamyanba, Jung, Soon Ki, Hammond, Tracy
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916422897106944
author Rahman, Md Maklachur
Tutul, Abdullah Aman
Nath, Ankur
Laishram, Lamyanba
Jung, Soon Ki
Hammond, Tracy
author_facet Rahman, Md Maklachur
Tutul, Abdullah Aman
Nath, Ankur
Laishram, Lamyanba
Jung, Soon Ki
Hammond, Tracy
contents Mamba is emerging as a novel approach to overcome the challenges faced by Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) in computer vision. While CNNs excel at extracting local features, they often struggle to capture long-range dependencies without complex architectural modifications. In contrast, ViTs effectively model global relationships but suffer from high computational costs due to the quadratic complexity of their self-attention mechanisms. Mamba addresses these limitations by leveraging Selective Structured State Space Models to effectively capture long-range dependencies with linear computational complexity. This survey analyzes the unique contributions, computational benefits, and applications of Mamba models while also identifying challenges and potential future research directions. We provide a foundational resource for advancing the understanding and growth of Mamba models in computer vision. An overview of this work is available at https://github.com/maklachur/Mamba-in-Computer-Vision.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03105
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mamba in Vision: A Comprehensive Survey of Techniques and Applications
Rahman, Md Maklachur
Tutul, Abdullah Aman
Nath, Ankur
Laishram, Lamyanba
Jung, Soon Ki
Hammond, Tracy
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Mamba is emerging as a novel approach to overcome the challenges faced by Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) in computer vision. While CNNs excel at extracting local features, they often struggle to capture long-range dependencies without complex architectural modifications. In contrast, ViTs effectively model global relationships but suffer from high computational costs due to the quadratic complexity of their self-attention mechanisms. Mamba addresses these limitations by leveraging Selective Structured State Space Models to effectively capture long-range dependencies with linear computational complexity. This survey analyzes the unique contributions, computational benefits, and applications of Mamba models while also identifying challenges and potential future research directions. We provide a foundational resource for advancing the understanding and growth of Mamba models in computer vision. An overview of this work is available at https://github.com/maklachur/Mamba-in-Computer-Vision.
title Mamba in Vision: A Comprehensive Survey of Techniques and Applications
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.03105