PackUV: Packed Gaussian UV Maps for 4D Volumetric Video

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rai, Aashish, Xing, Angela, Agarwal, Anushka, Cong, Xiaoyan, Li, Zekun, Lu, Tao, Prakash, Aayush, Sridhar, Srinath
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911495034503168
author Rai, Aashish
Xing, Angela
Agarwal, Anushka
Cong, Xiaoyan
Li, Zekun
Lu, Tao
Prakash, Aayush
Sridhar, Srinath
author_facet Rai, Aashish
Xing, Angela
Agarwal, Anushka
Cong, Xiaoyan
Li, Zekun
Lu, Tao
Prakash, Aayush
Sridhar, Srinath
contents Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality reconstruction but break down on long sequences, temporal inconsistency, and fail under large motions and disocclusions. Moreover, their outputs are typically incompatible with conventional video coding pipelines, preventing practical applications. We introduce PackUV, a novel 4D Gaussian representation that maps all Gaussian attributes into a sequence of structured, multi-scale UV atlas, enabling compact, image-native storage. To fit this representation from multi-view videos, we propose PackUV-GS, a temporally consistent fitting method that directly optimizes Gaussian parameters in the UV domain. A flow-guided Gaussian labeling and video keyframing module identifies dynamic Gaussians, stabilizes static regions, and preserves temporal coherence even under large motions and disocclusions. The resulting UV atlas format is the first unified volumetric video representation compatible with standard video codecs (e.g., FFV1) without losing quality, enabling efficient streaming within existing multimedia infrastructure. To evaluate long-duration volumetric capture, we present PackUV-2B, the largest multi-view video dataset to date, featuring more than 50 synchronized cameras, substantial motion, and frequent disocclusions across 100 sequences and 2B (billion) frames. Extensive experiments demonstrate that our method surpasses existing baselines in rendering fidelity while scaling to sequences up to 30 minutes with consistent quality.
format Preprint
id arxiv_https___arxiv_org_abs_2602_23040
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
Rai, Aashish
Xing, Angela
Agarwal, Anushka
Cong, Xiaoyan
Li, Zekun
Lu, Tao
Prakash, Aayush
Sridhar, Srinath
Computer Vision and Pattern Recognition
Volumetric videos offer immersive 4D experiences, but remain difficult to reconstruct, store, and stream at scale. Existing Gaussian Splatting based methods achieve high-quality reconstruction but break down on long sequences, temporal inconsistency, and fail under large motions and disocclusions. Moreover, their outputs are typically incompatible with conventional video coding pipelines, preventing practical applications. We introduce PackUV, a novel 4D Gaussian representation that maps all Gaussian attributes into a sequence of structured, multi-scale UV atlas, enabling compact, image-native storage. To fit this representation from multi-view videos, we propose PackUV-GS, a temporally consistent fitting method that directly optimizes Gaussian parameters in the UV domain. A flow-guided Gaussian labeling and video keyframing module identifies dynamic Gaussians, stabilizes static regions, and preserves temporal coherence even under large motions and disocclusions. The resulting UV atlas format is the first unified volumetric video representation compatible with standard video codecs (e.g., FFV1) without losing quality, enabling efficient streaming within existing multimedia infrastructure. To evaluate long-duration volumetric capture, we present PackUV-2B, the largest multi-view video dataset to date, featuring more than 50 synchronized cameras, substantial motion, and frequent disocclusions across 100 sequences and 2B (billion) frames. Extensive experiments demonstrate that our method surpasses existing baselines in rendering fidelity while scaling to sequences up to 30 minutes with consistent quality.
title PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.23040