GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Longan, Shi, Yuang, Ooi, Wei Tsang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910794082418688
author Wang, Longan
Shi, Yuang
Ooi, Wei Tsang
author_facet Wang, Longan
Shi, Yuang
Ooi, Wei Tsang
contents 3D Gaussian splats have emerged as a revolutionary, effective, learned representation for static 3D scenes. In this work, we explore using 2D Gaussian splats as a new primitive for representing videos. We propose GSVC, an approach to learning a set of 2D Gaussian splats that can effectively represent and compress video frames. GSVC incorporates the following techniques: (i) To exploit temporal redundancy among adjacent frames, which can speed up training and improve the compression efficiency, we predict the Gaussian splats of a frame based on its previous frame; (ii) To control the trade-offs between file size and quality, we remove Gaussian splats with low contribution to the video quality; (iii) To capture dynamics in videos, we randomly add Gaussian splats to fit content with large motion or newly-appeared objects; (iv) To handle significant changes in the scene, we detect key frames based on loss differences during the learning process. Experiment results show that GSVC achieves good rate-distortion trade-offs, comparable to state-of-the-art video codecs such as AV1 and VVC, and a rendering speed of 1500 fps for a 1920x1080 video.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12060
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting
Wang, Longan
Shi, Yuang
Ooi, Wei Tsang
Computer Vision and Pattern Recognition
Multimedia
3D Gaussian splats have emerged as a revolutionary, effective, learned representation for static 3D scenes. In this work, we explore using 2D Gaussian splats as a new primitive for representing videos. We propose GSVC, an approach to learning a set of 2D Gaussian splats that can effectively represent and compress video frames. GSVC incorporates the following techniques: (i) To exploit temporal redundancy among adjacent frames, which can speed up training and improve the compression efficiency, we predict the Gaussian splats of a frame based on its previous frame; (ii) To control the trade-offs between file size and quality, we remove Gaussian splats with low contribution to the video quality; (iii) To capture dynamics in videos, we randomly add Gaussian splats to fit content with large motion or newly-appeared objects; (iv) To handle significant changes in the scene, we detect key frames based on loss differences during the learning process. Experiment results show that GSVC achieves good rate-distortion trade-offs, comparable to state-of-the-art video codecs such as AV1 and VVC, and a rendering speed of 1500 fps for a 1920x1080 video.
title GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2501.12060