Diffusion Masked Pretraining for Dynamic Point Cloud

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhuoyue, Zhu, Jihua, Fang, Chaowei, Liu, Jian, Mian, Ajmal Saeed
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914550934142976
author Zhang, Zhuoyue
Zhu, Jihua
Fang, Chaowei
Liu, Jian
Mian, Ajmal Saeed
author_facet Zhang, Zhuoyue
Zhu, Jihua
Fang, Chaowei
Liu, Jian
Mian, Ajmal Saeed
contents Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground-truth tube centers as decoder positional embeddings, causing spatio-temporal positional leakage. Moreover, they supervise inter-frame motion with deterministic proxy targets that systematically discard distributional structure by collapsing multimodal trajectory uncertainty into conditional means. To address these limitations, we propose Diffusion Masked Pretraining (DiMP), a unified self-supervised framework for dynamic point clouds. DiMP introduces diffusion modeling into both positional inference and motion learning. It first applies forward diffusion noise only to masked tube centers, then predicts clean centers from visible spatio-temporal context. This removes positional leakage while preserving visible coordinates as clean temporal anchors. DiMP also reformulates point-wise inter-frame displacement supervision as a DDPM noise-prediction objective conditioned on decoded representations. This design drives the encoder to target the full conditional distribution of plausible motions under a variational surrogate, rather than collapsing to a single deterministic estimate. Extensive experiments demonstrate that DiMP consistently improves downstream accuracy over the backbone alone, with absolute gains of 11.21% on offline action segmentation and 13.65% under causally constrained online inference.Codes are available at https://github.com/InitalZ/DiMP.git.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03639
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Diffusion Masked Pretraining for Dynamic Point Cloud
Zhang, Zhuoyue
Zhu, Jihua
Fang, Chaowei
Liu, Jian
Mian, Ajmal Saeed
Computer Vision and Pattern Recognition
Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground-truth tube centers as decoder positional embeddings, causing spatio-temporal positional leakage. Moreover, they supervise inter-frame motion with deterministic proxy targets that systematically discard distributional structure by collapsing multimodal trajectory uncertainty into conditional means. To address these limitations, we propose Diffusion Masked Pretraining (DiMP), a unified self-supervised framework for dynamic point clouds. DiMP introduces diffusion modeling into both positional inference and motion learning. It first applies forward diffusion noise only to masked tube centers, then predicts clean centers from visible spatio-temporal context. This removes positional leakage while preserving visible coordinates as clean temporal anchors. DiMP also reformulates point-wise inter-frame displacement supervision as a DDPM noise-prediction objective conditioned on decoded representations. This design drives the encoder to target the full conditional distribution of plausible motions under a variational surrogate, rather than collapsing to a single deterministic estimate. Extensive experiments demonstrate that DiMP consistently improves downstream accuracy over the backbone alone, with absolute gains of 11.21% on offline action segmentation and 13.65% under causally constrained online inference.Codes are available at https://github.com/InitalZ/DiMP.git.
title Diffusion Masked Pretraining for Dynamic Point Cloud
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.03639