Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liang, Hanxue, Ren, Jiawei, Mirzaei, Ashkan, Torralba, Antonio, Liu, Ziwei, Gilitschenski, Igor, Fidler, Sanja, Oztireli, Cengiz, Ling, Huan, Gojcic, Zan, Huang, Jiahui
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916959449251840
author Liang, Hanxue
Ren, Jiawei
Mirzaei, Ashkan
Torralba, Antonio
Liu, Ziwei
Gilitschenski, Igor
Fidler, Sanja
Oztireli, Cengiz
Ling, Huan
Gojcic, Zan
Huang, Jiahui
author_facet Liang, Hanxue
Ren, Jiawei
Mirzaei, Ashkan
Torralba, Antonio
Liu, Ziwei
Gilitschenski, Igor
Fidler, Sanja
Oztireli, Cengiz
Ling, Huan
Gojcic, Zan
Huang, Jiahui
contents Recent advancements in static feed-forward scene reconstruction have demonstrated significant progress in high-quality novel view synthesis. However, these models often struggle with generalizability across diverse environments and fail to effectively handle dynamic content. We present BTimer (short for BulletTimer), the first motion-aware feed-forward model for real-time reconstruction and novel view synthesis of dynamic scenes. Our approach reconstructs the full scene in a 3D Gaussian Splatting representation at a given target ('bullet') timestamp by aggregating information from all the context frames. Such a formulation allows BTimer to gain scalability and generalization by leveraging both static and dynamic scene datasets. Given a casual monocular dynamic video, BTimer reconstructs a bullet-time scene within 150ms while reaching state-of-the-art performance on both static and dynamic scene datasets, even compared with optimization-based approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03526
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos
Liang, Hanxue
Ren, Jiawei
Mirzaei, Ashkan
Torralba, Antonio
Liu, Ziwei
Gilitschenski, Igor
Fidler, Sanja
Oztireli, Cengiz
Ling, Huan
Gojcic, Zan
Huang, Jiahui
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Recent advancements in static feed-forward scene reconstruction have demonstrated significant progress in high-quality novel view synthesis. However, these models often struggle with generalizability across diverse environments and fail to effectively handle dynamic content. We present BTimer (short for BulletTimer), the first motion-aware feed-forward model for real-time reconstruction and novel view synthesis of dynamic scenes. Our approach reconstructs the full scene in a 3D Gaussian Splatting representation at a given target ('bullet') timestamp by aggregating information from all the context frames. Such a formulation allows BTimer to gain scalability and generalization by leveraging both static and dynamic scene datasets. Given a casual monocular dynamic video, BTimer reconstructs a bullet-time scene within 150ms while reaching state-of-the-art performance on both static and dynamic scene datasets, even compared with optimization-based approaches.
title Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
url https://arxiv.org/abs/2412.03526