SS4D: Native 4D Generative Model via Structured Spacetime Latents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhibing, Zhang, Mengchen, Wu, Tong, Tan, Jing, Wang, Jiaqi, Lin, Dahua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914203900575744
author Li, Zhibing
Zhang, Mengchen
Wu, Tong
Tan, Jing
Wang, Jiaqi
Lin, Dahua
author_facet Li, Zhibing
Zhang, Mengchen
Wu, Tong
Tan, Jing
Wang, Jiaqi
Lin, Dahua
contents We present SS4D, a native 4D generative model that synthesizes dynamic 3D objects directly from monocular video. Unlike prior approaches that construct 4D representations by optimizing over 3D or video generative models, we train a generator directly on 4D data, achieving high fidelity, temporal coherence, and structural consistency. At the core of our method is a compressed set of structured spacetime latents. Specifically, (1) To address the scarcity of 4D training data, we build on a pre-trained single-image-to-3D model, preserving strong spatial consistency. (2) Temporal consistency is enforced by introducing dedicated temporal layers that reason across frames. (3) To support efficient training and inference over long video sequences, we compress the latent sequence along the temporal axis using factorized 4D convolutions and temporal downsampling blocks. In addition, we employ a carefully designed training strategy to enhance robustness against occlusion
format Preprint
id arxiv_https___arxiv_org_abs_2512_14284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SS4D: Native 4D Generative Model via Structured Spacetime Latents
Li, Zhibing
Zhang, Mengchen
Wu, Tong
Tan, Jing
Wang, Jiaqi
Lin, Dahua
Computer Vision and Pattern Recognition
We present SS4D, a native 4D generative model that synthesizes dynamic 3D objects directly from monocular video. Unlike prior approaches that construct 4D representations by optimizing over 3D or video generative models, we train a generator directly on 4D data, achieving high fidelity, temporal coherence, and structural consistency. At the core of our method is a compressed set of structured spacetime latents. Specifically, (1) To address the scarcity of 4D training data, we build on a pre-trained single-image-to-3D model, preserving strong spatial consistency. (2) Temporal consistency is enforced by introducing dedicated temporal layers that reason across frames. (3) To support efficient training and inference over long video sequences, we compress the latent sequence along the temporal axis using factorized 4D convolutions and temporal downsampling blocks. In addition, we employ a carefully designed training strategy to enhance robustness against occlusion
title SS4D: Native 4D Generative Model via Structured Spacetime Latents
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.14284