Saved in:
Bibliographic Details
Main Authors: Lin, Xuanyu, Zeng, Xiaona, Zheng, Xianwei, Li, Xutao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.17296
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918102649798656
author Lin, Xuanyu
Zeng, Xiaona
Zheng, Xianwei
Li, Xutao
author_facet Lin, Xuanyu
Zeng, Xiaona
Zheng, Xianwei
Li, Xutao
contents Mamba has recently gained widespread attention as a backbone model for point cloud modeling, leveraging a state-space architecture that enables efficient global sequence modeling with linear complexity. However, its lack of local inductive bias limits its capacity to capture fine-grained geometric structures in 3D data. To address this limitation, we propose \textbf{PointLAMA}, a point cloud pretraining framework that combines task-aware point cloud serialization, a hybrid encoder with integrated Latent Attention and Mamba blocks, and a conditional diffusion mechanism built upon the Mamba backbone. Specifically, the task-aware point cloud serialization employs Hilbert/Trans-Hilbert space-filling curves and axis-wise sorting to structurally align point tokens for classification and segmentation tasks, respectively. Our lightweight Latent Attention block features a Point-wise Multi-head Latent Attention (PMLA) module, which is specifically designed to align with the Mamba architecture by leveraging the shared latent space characteristics of PMLA and Mamba. This enables enhanced local context modeling while preserving overall efficiency. To further enhance representation learning, we incorporate a conditional diffusion mechanism during pretraining, which denoises perturbed feature sequences without relying on explicit point-wise reconstruction. Experimental results demonstrate that PointLAMA achieves competitive performance on multiple benchmark datasets with minimal parameter count and FLOPs, validating its effectiveness for efficient point cloud pretraining.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17296
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PointLAMA: Latent Attention meets Mamba for Efficient Point Cloud Pretraining
Lin, Xuanyu
Zeng, Xiaona
Zheng, Xianwei
Li, Xutao
Computer Vision and Pattern Recognition
Mamba has recently gained widespread attention as a backbone model for point cloud modeling, leveraging a state-space architecture that enables efficient global sequence modeling with linear complexity. However, its lack of local inductive bias limits its capacity to capture fine-grained geometric structures in 3D data. To address this limitation, we propose \textbf{PointLAMA}, a point cloud pretraining framework that combines task-aware point cloud serialization, a hybrid encoder with integrated Latent Attention and Mamba blocks, and a conditional diffusion mechanism built upon the Mamba backbone. Specifically, the task-aware point cloud serialization employs Hilbert/Trans-Hilbert space-filling curves and axis-wise sorting to structurally align point tokens for classification and segmentation tasks, respectively. Our lightweight Latent Attention block features a Point-wise Multi-head Latent Attention (PMLA) module, which is specifically designed to align with the Mamba architecture by leveraging the shared latent space characteristics of PMLA and Mamba. This enables enhanced local context modeling while preserving overall efficiency. To further enhance representation learning, we incorporate a conditional diffusion mechanism during pretraining, which denoises perturbed feature sequences without relying on explicit point-wise reconstruction. Experimental results demonstrate that PointLAMA achieves competitive performance on multiple benchmark datasets with minimal parameter count and FLOPs, validating its effectiveness for efficient point cloud pretraining.
title PointLAMA: Latent Attention meets Mamba for Efficient Point Cloud Pretraining
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.17296