DyST: Towards Dynamic Neural Scene Representations on Real-World Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Seitzer, Maximilian, van Steenkiste, Sjoerd, Kipf, Thomas, Greff, Klaus, Sajjadi, Mehdi S. M.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909137507450880
author Seitzer, Maximilian
van Steenkiste, Sjoerd
Kipf, Thomas
Greff, Klaus
Sajjadi, Mehdi S. M.
author_facet Seitzer, Maximilian
van Steenkiste, Sjoerd
Kipf, Thomas
Greff, Klaus
Sajjadi, Mehdi S. M.
contents Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world videos. Our Dynamic Scene Transformer (DyST) model leverages recent work in neural scene representation to learn a latent decomposition of monocular real-world videos into scene content, per-view scene dynamics, and camera pose. This separation is achieved through a novel co-training scheme on monocular videos and our new synthetic dataset DySO. DyST learns tangible latent representations for dynamic scenes that enable view generation with separate control over the camera and the content of the scene.
format Preprint
id arxiv_https___arxiv_org_abs_2310_06020
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
Seitzer, Maximilian
van Steenkiste, Sjoerd
Kipf, Thomas
Greff, Klaus
Sajjadi, Mehdi S. M.
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Robotics
Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world videos. Our Dynamic Scene Transformer (DyST) model leverages recent work in neural scene representation to learn a latent decomposition of monocular real-world videos into scene content, per-view scene dynamics, and camera pose. This separation is achieved through a novel co-training scheme on monocular videos and our new synthetic dataset DySO. DyST learns tangible latent representations for dynamic scenes that enable view generation with separate control over the camera and the content of the scene.
title DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Robotics
url https://arxiv.org/abs/2310.06020