VeGaS: Video Gaussian Splatting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Smolak-Dyżewska, Weronika, Malarz, Dawid, Howil, Kornel, Kaczmarczyk, Jan, Mazur, Marcin, Spurek, Przemysław
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912910722203648
author Smolak-Dyżewska, Weronika
Malarz, Dawid
Howil, Kornel
Kaczmarczyk, Jan
Mazur, Marcin
Spurek, Przemysław
author_facet Smolak-Dyżewska, Weronika
Malarz, Dawid
Howil, Kornel
Kaczmarczyk, Jan
Mazur, Marcin
Spurek, Przemysław
contents Implicit Neural Representations (INRs) employ neural networks to approximate discrete data as continuous functions. In the context of video data, such models can be utilized to transform the coordinates of pixel locations along with frame occurrence times (or indices) into RGB color values. Although INRs facilitate effective compression, they are unsuitable for editing purposes. One potential solution is to use a 3D Gaussian Splatting (3DGS) based model, such as the Video Gaussian Representation (VGR), which is capable of encoding video as a multitude of 3D Gaussians and is applicable for numerous video processing operations, including editing. Nevertheless, in this case, the capacity for modification is constrained to a limited set of basic transformations. To address this issue, we introduce the Video Gaussian Splatting (VeGaS) model, which enables realistic modifications of video data. To construct VeGaS, we propose a novel family of Folded-Gaussian distributions designed to capture nonlinear dynamics in a video stream and model consecutive frames by 2D Gaussians obtained as respective conditional distributions. Our experiments demonstrate that VeGaS outperforms state-of-the-art solutions in frame reconstruction tasks and allows realistic modifications of video data. The code is available at: https://github.com/gmum/VeGaS.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11024
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VeGaS: Video Gaussian Splatting
Smolak-Dyżewska, Weronika
Malarz, Dawid
Howil, Kornel
Kaczmarczyk, Jan
Mazur, Marcin
Spurek, Przemysław
Computer Vision and Pattern Recognition
Implicit Neural Representations (INRs) employ neural networks to approximate discrete data as continuous functions. In the context of video data, such models can be utilized to transform the coordinates of pixel locations along with frame occurrence times (or indices) into RGB color values. Although INRs facilitate effective compression, they are unsuitable for editing purposes. One potential solution is to use a 3D Gaussian Splatting (3DGS) based model, such as the Video Gaussian Representation (VGR), which is capable of encoding video as a multitude of 3D Gaussians and is applicable for numerous video processing operations, including editing. Nevertheless, in this case, the capacity for modification is constrained to a limited set of basic transformations. To address this issue, we introduce the Video Gaussian Splatting (VeGaS) model, which enables realistic modifications of video data. To construct VeGaS, we propose a novel family of Folded-Gaussian distributions designed to capture nonlinear dynamics in a video stream and model consecutive frames by 2D Gaussians obtained as respective conditional distributions. Our experiments demonstrate that VeGaS outperforms state-of-the-art solutions in frame reconstruction tasks and allows realistic modifications of video data. The code is available at: https://github.com/gmum/VeGaS.
title VeGaS: Video Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.11024