Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yoon, Dongho, Lee, Gungyu, Chang, Jaewon, Lee, Yunjae, Lee, Dongjae, Rhu, Minsoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912519809925120
author Yoon, Dongho
Lee, Gungyu
Chang, Jaewon
Lee, Yunjae
Lee, Dongjae
Rhu, Minsoo
author_facet Yoon, Dongho
Lee, Gungyu
Chang, Jaewon
Lee, Yunjae
Lee, Dongjae
Rhu, Minsoo
contents Transformers have proven effective in language modeling but are limited by high computational and memory demands that grow quadratically with input sequence length. State space models (SSMs) offer a promising alternative by reducing attention complexity from $O(L^2)$ to $O(L)$ while also lowering overall memory consumption. Vision Mamba adapts the SSM approach for computer vision tasks, achieving lower latency and memory consumption than traditional transformer models. However, deploying Vision Mamba on edge devices is challenging due to its sequential scan operations, which hinder GPU efficiency. We propose Mamba-X, an end-to-end Vision Mamba accelerator that includes a systolic scan array to maximize parallelism and minimize memory traffic, along with a hybrid, hardware-friendly quantization technique to reduce memory usage and improve hardware efficiency without sacrificing accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02977
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
Yoon, Dongho
Lee, Gungyu
Chang, Jaewon
Lee, Yunjae
Lee, Dongjae
Rhu, Minsoo
Hardware Architecture
Transformers have proven effective in language modeling but are limited by high computational and memory demands that grow quadratically with input sequence length. State space models (SSMs) offer a promising alternative by reducing attention complexity from $O(L^2)$ to $O(L)$ while also lowering overall memory consumption. Vision Mamba adapts the SSM approach for computer vision tasks, achieving lower latency and memory consumption than traditional transformer models. However, deploying Vision Mamba on edge devices is challenging due to its sequential scan operations, which hinder GPU efficiency. We propose Mamba-X, an end-to-end Vision Mamba accelerator that includes a systolic scan array to maximize parallelism and minimize memory traffic, along with a hybrid, hardware-friendly quantization technique to reduce memory usage and improve hardware efficiency without sacrificing accuracy.
title Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
topic Hardware Architecture
url https://arxiv.org/abs/2508.02977