Splatter a Video: Video Gaussian Representation for Versatile Processing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Yang-Tian, Huang, Yi-Hua, Ma, Lin, Lyu, Xiaoyang, Cao, Yan-Pei, Qi, Xiaojuan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911933280550912
author Sun, Yang-Tian
Huang, Yi-Hua
Ma, Lin
Lyu, Xiaoyang
Cao, Yan-Pei
Qi, Xiaojuan
author_facet Sun, Yang-Tian
Huang, Yi-Hua
Ma, Lin
Lyu, Xiaoyang
Cao, Yan-Pei
Qi, Xiaojuan
contents Video representation is a long-standing problem that is crucial for various down-stream tasks, such as tracking,depth prediction,segmentation,view synthesis,and editing. However, current methods either struggle to model complex motions due to the absence of 3D structure or rely on implicit 3D representations that are ill-suited for manipulation tasks. To address these challenges, we introduce a novel explicit 3D representation-video Gaussian representation -- that embeds a video into 3D Gaussians. Our proposed representation models video appearance in a 3D canonical space using explicit Gaussians as proxies and associates each Gaussian with 3D motions for video motion. This approach offers a more intrinsic and explicit representation than layered atlas or volumetric pixel matrices. To obtain such a representation, we distill 2D priors, such as optical flow and depth, from foundation models to regularize learning in this ill-posed setting. Extensive applications demonstrate the versatility of our new video representation. It has been proven effective in numerous video processing tasks, including tracking, consistent video depth and feature refinement, motion and appearance editing, and stereoscopic video generation. Project page: https://sunyangtian.github.io/spatter_a_video_web/
format Preprint
id arxiv_https___arxiv_org_abs_2406_13870
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Splatter a Video: Video Gaussian Representation for Versatile Processing
Sun, Yang-Tian
Huang, Yi-Hua
Ma, Lin
Lyu, Xiaoyang
Cao, Yan-Pei
Qi, Xiaojuan
Computer Vision and Pattern Recognition
Video representation is a long-standing problem that is crucial for various down-stream tasks, such as tracking,depth prediction,segmentation,view synthesis,and editing. However, current methods either struggle to model complex motions due to the absence of 3D structure or rely on implicit 3D representations that are ill-suited for manipulation tasks. To address these challenges, we introduce a novel explicit 3D representation-video Gaussian representation -- that embeds a video into 3D Gaussians. Our proposed representation models video appearance in a 3D canonical space using explicit Gaussians as proxies and associates each Gaussian with 3D motions for video motion. This approach offers a more intrinsic and explicit representation than layered atlas or volumetric pixel matrices. To obtain such a representation, we distill 2D priors, such as optical flow and depth, from foundation models to regularize learning in this ill-posed setting. Extensive applications demonstrate the versatility of our new video representation. It has been proven effective in numerous video processing tasks, including tracking, consistent video depth and feature refinement, motion and appearance editing, and stereoscopic video generation. Project page: https://sunyangtian.github.io/spatter_a_video_web/
title Splatter a Video: Video Gaussian Representation for Versatile Processing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.13870