Characterizing the Behavior of Training Mamba-based State Space Models on GPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baruah, Trinayan, Shivdikar, Kaustubh, Prescott, Sara, Kaeli, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911119799484416
author Baruah, Trinayan
Shivdikar, Kaustubh
Prescott, Sara
Kaeli, David
author_facet Baruah, Trinayan
Shivdikar, Kaustubh
Prescott, Sara
Kaeli, David
contents Mamba-based State Space Models (SSM) have emerged as a promising alternative to the ubiquitous transformers. Despite the expressive power of transformers, the quadratic complexity of computing attention is a major impediment to scaling performance as we increase the sequence length. SSMs provide an alternative path that addresses this problem, reducing the computational complexity requirements of self-attention with novel model architectures for different domains and fields such as video, text generation and graphs. Thus, it is important to characterize the behavior of these emerging workloads on GPUs and understand their requirements during GPU microarchitectural design. In this work we evaluate Mamba-based SSMs and characterize their behavior during training on GPUs. We construct a workload suite that offers representative models that span different model architectures. We then use this suite to analyze the architectural implications of running Mamba-based SSMs on GPUs. Our work sheds new light on potential optimizations to continue scaling the performance for such models.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17679
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
Baruah, Trinayan
Shivdikar, Kaustubh
Prescott, Sara
Kaeli, David
Machine Learning
Hardware Architecture
Computation and Language
Mamba-based State Space Models (SSM) have emerged as a promising alternative to the ubiquitous transformers. Despite the expressive power of transformers, the quadratic complexity of computing attention is a major impediment to scaling performance as we increase the sequence length. SSMs provide an alternative path that addresses this problem, reducing the computational complexity requirements of self-attention with novel model architectures for different domains and fields such as video, text generation and graphs. Thus, it is important to characterize the behavior of these emerging workloads on GPUs and understand their requirements during GPU microarchitectural design. In this work we evaluate Mamba-based SSMs and characterize their behavior during training on GPUs. We construct a workload suite that offers representative models that span different model architectures. We then use this suite to analyze the architectural implications of running Mamba-based SSMs on GPUs. Our work sheds new light on potential optimizations to continue scaling the performance for such models.
title Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
topic Machine Learning
Hardware Architecture
Computation and Language
url https://arxiv.org/abs/2508.17679