STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wenhao, Jiang, Xueying, Zhang, Gongjie, Zhang, Xiaoqin, Shao, Ling, Lu, Shijian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911588890443776
author Li, Wenhao
Jiang, Xueying
Zhang, Gongjie
Zhang, Xiaoqin
Shao, Ling
Lu, Shijian
author_facet Li, Wenhao
Jiang, Xueying
Zhang, Gongjie
Zhang, Xiaoqin
Shao, Ling
Lu, Shijian
contents 4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the spatiotemporal domain where the underlying geometric characteristics of 4D point cloud videos are hard to capture, leading to degraded representation learning and understanding of 4D point cloud videos. We address the above challenge from a complementary spectral perspective. By transforming 4D point cloud videos into graph spectral signals, we can decompose them into multiple frequency bands each of which captures distinct geometric structures of point cloud videos. Our spectral analysis reveals that the decomposed low-frequency signals capture more coarse shapes while high-frequency signals encode more fine-grained geometry details. Building on these observations, we design Spatio-Temporal-Spectral Mixer (STS-Mixer), a unified framework that mixes spatial, temporal, and spectral representations of point cloud videos. STS-Mixer integrates multi-band delineated spectral signals with spatiotemporal information to capture rich geometries and temporal dynamics, while enabling fine-grained and holistic understanding of 4D point cloud videos. Extensive experiments show that STS-Mixer achieves superior performance consistently across multiple widely adopted benchmarks on both 3D action recognition and 4D semantic segmentation tasks. Code and models are available at https://github.com/Vegetebird/STS-Mixer.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11637
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
Li, Wenhao
Jiang, Xueying
Zhang, Gongjie
Zhang, Xiaoqin
Shao, Ling
Lu, Shijian
Computer Vision and Pattern Recognition
4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the spatiotemporal domain where the underlying geometric characteristics of 4D point cloud videos are hard to capture, leading to degraded representation learning and understanding of 4D point cloud videos. We address the above challenge from a complementary spectral perspective. By transforming 4D point cloud videos into graph spectral signals, we can decompose them into multiple frequency bands each of which captures distinct geometric structures of point cloud videos. Our spectral analysis reveals that the decomposed low-frequency signals capture more coarse shapes while high-frequency signals encode more fine-grained geometry details. Building on these observations, we design Spatio-Temporal-Spectral Mixer (STS-Mixer), a unified framework that mixes spatial, temporal, and spectral representations of point cloud videos. STS-Mixer integrates multi-band delineated spectral signals with spatiotemporal information to capture rich geometries and temporal dynamics, while enabling fine-grained and holistic understanding of 4D point cloud videos. Extensive experiments show that STS-Mixer achieves superior performance consistently across multiple widely adopted benchmarks on both 3D action recognition and 4D semantic segmentation tasks. Code and models are available at https://github.com/Vegetebird/STS-Mixer.
title STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.11637