PointMamba: A Simple State Space Model for Point Cloud Analysis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liang, Dingkang, Zhou, Xin, Xu, Wei, Zhu, Xingkui, Zou, Zhikang, Ye, Xiaoqing, Tan, Xiao, Bai, Xiang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909401941540864
author Liang, Dingkang
Zhou, Xin
Xu, Wei
Zhu, Xingkui
Zou, Zhikang
Ye, Xiaoqing
Tan, Xiao
Bai, Xiang
author_facet Liang, Dingkang
Zhou, Xin
Xu, Wei
Zhu, Xingkui
Zou, Zhikang
Ye, Xiaoqing
Tan, Xiao
Bai, Xiang
contents Transformers have become one of the foundational architectures in point cloud analysis tasks due to their excellent global modeling ability. However, the attention mechanism has quadratic complexity, making the design of a linear complexity method with global modeling appealing. In this paper, we propose PointMamba, transferring the success of Mamba, a recent representative state space model (SSM), from NLP to point cloud analysis tasks. Unlike traditional Transformers, PointMamba employs a linear complexity algorithm, presenting global modeling capacity while significantly reducing computational costs. Specifically, our method leverages space-filling curves for effective point tokenization and adopts an extremely simple, non-hierarchical Mamba encoder as the backbone. Comprehensive evaluations demonstrate that PointMamba achieves superior performance across multiple datasets while significantly reducing GPU memory usage and FLOPs. This work underscores the potential of SSMs in 3D vision-related tasks and presents a simple yet effective Mamba-based baseline for future research. The code will be made available at \url{https://github.com/LMD0311/PointMamba}.
format Preprint
id arxiv_https___arxiv_org_abs_2402_10739
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PointMamba: A Simple State Space Model for Point Cloud Analysis
Liang, Dingkang
Zhou, Xin
Xu, Wei
Zhu, Xingkui
Zou, Zhikang
Ye, Xiaoqing
Tan, Xiao
Bai, Xiang
Computer Vision and Pattern Recognition
Transformers have become one of the foundational architectures in point cloud analysis tasks due to their excellent global modeling ability. However, the attention mechanism has quadratic complexity, making the design of a linear complexity method with global modeling appealing. In this paper, we propose PointMamba, transferring the success of Mamba, a recent representative state space model (SSM), from NLP to point cloud analysis tasks. Unlike traditional Transformers, PointMamba employs a linear complexity algorithm, presenting global modeling capacity while significantly reducing computational costs. Specifically, our method leverages space-filling curves for effective point tokenization and adopts an extremely simple, non-hierarchical Mamba encoder as the backbone. Comprehensive evaluations demonstrate that PointMamba achieves superior performance across multiple datasets while significantly reducing GPU memory usage and FLOPs. This work underscores the potential of SSMs in 3D vision-related tasks and presents a simple yet effective Mamba-based baseline for future research. The code will be made available at \url{https://github.com/LMD0311/PointMamba}.
title PointMamba: A Simple State Space Model for Point Cloud Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.10739