PointTransformerX: Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Reichardt, Laurenz, Ebert, Nikolas, Wasenmüller, Oliver
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917446790676480
author Reichardt, Laurenz
Ebert, Nikolas
Wasenmüller, Oliver
author_facet Reichardt, Laurenz
Ebert, Nikolas
Wasenmüller, Oliver
contents 3D point cloud perception remains tightly coupled to custom CUDA operators for spatial operations, limiting portability and efficiency on non-NVIDIA, AMD, and embedded hardware. We introduce PointTransformerX (PTX), a fully PyTorch-native vision transformer backbone for 3D point clouds, removing all custom CUDA operators and external libraries while retaining competitive accuracy. PTX introduces 3D-GS-RoPE, a rotary positional embedding that encodes 3D spatial relationships directly in self-attention without neighborhood construction, and further replaces sparse convolutional patch embedding with a linear projection. PTX explores inference-time scaling of attention windows to improve accuracy without retraining. With a redesigned feed-forward network, PTX achieves 98.7\% of PointTransformer V3's accuracy on ScanNet with 79.2\% fewer parameters and executing 1.6\times faster while requiring just 253 MB memory. PTX runs natively on NVIDIA GPUs, AMD GPUs (ROCm), and CPUs, providing an efficient and portable foundation for point cloud perception.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24169
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PointTransformerX: Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms
Reichardt, Laurenz
Ebert, Nikolas
Wasenmüller, Oliver
Computer Vision and Pattern Recognition
3D point cloud perception remains tightly coupled to custom CUDA operators for spatial operations, limiting portability and efficiency on non-NVIDIA, AMD, and embedded hardware. We introduce PointTransformerX (PTX), a fully PyTorch-native vision transformer backbone for 3D point clouds, removing all custom CUDA operators and external libraries while retaining competitive accuracy. PTX introduces 3D-GS-RoPE, a rotary positional embedding that encodes 3D spatial relationships directly in self-attention without neighborhood construction, and further replaces sparse convolutional patch embedding with a linear projection. PTX explores inference-time scaling of attention windows to improve accuracy without retraining. With a redesigned feed-forward network, PTX achieves 98.7\% of PointTransformer V3's accuracy on ScanNet with 79.2\% fewer parameters and executing 1.6\times faster while requiring just 253 MB memory. PTX runs natively on NVIDIA GPUs, AMD GPUs (ROCm), and CPUs, providing an efficient and portable foundation for point cloud perception.
title PointTransformerX: Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.24169