4-LEGS: 4D Language Embedded Gaussian Splatting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fiebelman, Gal, Cohen, Tamir, Morgenstern, Ayellet, Hedman, Peter, Averbuch-Elor, Hadar
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929713940791296
author Fiebelman, Gal
Cohen, Tamir
Morgenstern, Ayellet
Hedman, Peter
Averbuch-Elor, Hadar
author_facet Fiebelman, Gal
Cohen, Tamir
Morgenstern, Ayellet
Hedman, Peter
Averbuch-Elor, Hadar
contents The emergence of neural representations has revolutionized our means for digitally viewing a wide range of 3D scenes, enabling the synthesis of photorealistic images rendered from novel views. Recently, several techniques have been proposed for connecting these low-level representations with the high-level semantics understanding embodied within the scene. These methods elevate the rich semantic understanding from 2D imagery to 3D representations, distilling high-dimensional spatial features onto 3D space. In our work, we are interested in connecting language with a dynamic modeling of the world. We show how to lift spatio-temporal features to a 4D representation based on 3D Gaussian Splatting. This enables an interactive interface where the user can spatiotemporally localize events in the video from text prompts. We demonstrate our system on public 3D video datasets of people and animals performing various actions.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10719
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 4-LEGS: 4D Language Embedded Gaussian Splatting
Fiebelman, Gal
Cohen, Tamir
Morgenstern, Ayellet
Hedman, Peter
Averbuch-Elor, Hadar
Computer Vision and Pattern Recognition
Graphics
The emergence of neural representations has revolutionized our means for digitally viewing a wide range of 3D scenes, enabling the synthesis of photorealistic images rendered from novel views. Recently, several techniques have been proposed for connecting these low-level representations with the high-level semantics understanding embodied within the scene. These methods elevate the rich semantic understanding from 2D imagery to 3D representations, distilling high-dimensional spatial features onto 3D space. In our work, we are interested in connecting language with a dynamic modeling of the world. We show how to lift spatio-temporal features to a 4D representation based on 3D Gaussian Splatting. This enables an interactive interface where the user can spatiotemporally localize events in the video from text prompts. We demonstrate our system on public 3D video datasets of people and animals performing various actions.
title 4-LEGS: 4D Language Embedded Gaussian Splatting
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2410.10719