GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Inseo, Choi, Youngyoon, Lee, Joonseok
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910861491175424
author Lee, Inseo
Choi, Youngyoon
Lee, Joonseok
author_facet Lee, Inseo
Choi, Youngyoon
Lee, Joonseok
contents Implicit Neural Representation for Videos (NeRV) has introduced a novel paradigm for video representation and compression, outperforming traditional codecs. As model size grows, however, slow encoding and decoding speed and high memory consumption hinder its application in practice. To address these limitations, we propose a new video representation and compression method based on 2D Gaussian Splatting to efficiently handle video data. Our proposed deformable 2D Gaussian Splatting dynamically adapts the transformation of 2D Gaussians at each frame, significantly reducing memory cost. Equipped with a multi-plane-based spatiotemporal encoder and a lightweight decoder, it predicts changes in color, coordinates, and shape of initialized Gaussians, given the time step. By leveraging temporal gradients, our model effectively captures temporal redundancy at negligible cost, significantly enhancing video representation efficiency. Our method reduces GPU memory usage by up to 78.4%, and significantly expedites video processing, achieving 5.5x faster training and 12.5x faster decoding compared to the state-of-the-art NeRV methods.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
Lee, Inseo
Choi, Youngyoon
Lee, Joonseok
Computer Vision and Pattern Recognition
Implicit Neural Representation for Videos (NeRV) has introduced a novel paradigm for video representation and compression, outperforming traditional codecs. As model size grows, however, slow encoding and decoding speed and high memory consumption hinder its application in practice. To address these limitations, we propose a new video representation and compression method based on 2D Gaussian Splatting to efficiently handle video data. Our proposed deformable 2D Gaussian Splatting dynamically adapts the transformation of 2D Gaussians at each frame, significantly reducing memory cost. Equipped with a multi-plane-based spatiotemporal encoder and a lightweight decoder, it predicts changes in color, coordinates, and shape of initialized Gaussians, given the time step. By leveraging temporal gradients, our model effectively captures temporal redundancy at negligible cost, significantly enhancing video representation efficiency. Our method reduces GPU memory usage by up to 78.4%, and significantly expedites video processing, achieving 5.5x faster training and 12.5x faster decoding compared to the state-of-the-art NeRV methods.
title GaussianVideo: Efficient Video Representation and Compression by Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.04333