Mamba in Vision: A Comprehensive Survey of Techniques and Applications
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916422897106944 |
|---|---|
| author | Rahman, Md Maklachur Tutul, Abdullah Aman Nath, Ankur Laishram, Lamyanba Jung, Soon Ki Hammond, Tracy |
| author_facet | Rahman, Md Maklachur Tutul, Abdullah Aman Nath, Ankur Laishram, Lamyanba Jung, Soon Ki Hammond, Tracy |
| contents | Mamba is emerging as a novel approach to overcome the challenges faced by Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) in computer vision. While CNNs excel at extracting local features, they often struggle to capture long-range dependencies without complex architectural modifications. In contrast, ViTs effectively model global relationships but suffer from high computational costs due to the quadratic complexity of their self-attention mechanisms. Mamba addresses these limitations by leveraging Selective Structured State Space Models to effectively capture long-range dependencies with linear computational complexity. This survey analyzes the unique contributions, computational benefits, and applications of Mamba models while also identifying challenges and potential future research directions. We provide a foundational resource for advancing the understanding and growth of Mamba models in computer vision. An overview of this work is available at https://github.com/maklachur/Mamba-in-Computer-Vision. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_03105 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Mamba in Vision: A Comprehensive Survey of Techniques and Applications Rahman, Md Maklachur Tutul, Abdullah Aman Nath, Ankur Laishram, Lamyanba Jung, Soon Ki Hammond, Tracy Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning Mamba is emerging as a novel approach to overcome the challenges faced by Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) in computer vision. While CNNs excel at extracting local features, they often struggle to capture long-range dependencies without complex architectural modifications. In contrast, ViTs effectively model global relationships but suffer from high computational costs due to the quadratic complexity of their self-attention mechanisms. Mamba addresses these limitations by leveraging Selective Structured State Space Models to effectively capture long-range dependencies with linear computational complexity. This survey analyzes the unique contributions, computational benefits, and applications of Mamba models while also identifying challenges and potential future research directions. We provide a foundational resource for advancing the understanding and growth of Mamba models in computer vision. An overview of this work is available at https://github.com/maklachur/Mamba-in-Computer-Vision. |
| title | Mamba in Vision: A Comprehensive Survey of Techniques and Applications |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2410.03105 |