Efficiently Reconstructing Dynamic Scenes One D4RT at a Time

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Chuhan, Moing, Guillaume Le, Koppula, Skanda, Rocco, Ignacio, Momeni, Liliane, Xie, Junyu, Sun, Shuyang, Sukthankar, Rahul, Barral, Joëlle K., Hadsell, Raia, Ghahramani, Zoubin, Zisserman, Andrew, Zhang, Junlin, Sajjadi, Mehdi S. M.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911311191867392
author Zhang, Chuhan
Moing, Guillaume Le
Koppula, Skanda
Rocco, Ignacio
Momeni, Liliane
Xie, Junyu
Sun, Shuyang
Sukthankar, Rahul
Barral, Joëlle K.
Hadsell, Raia
Ghahramani, Zoubin
Zisserman, Andrew
Zhang, Junlin
Sajjadi, Mehdi S. M.
author_facet Zhang, Chuhan
Moing, Guillaume Le
Koppula, Skanda
Rocco, Ignacio
Momeni, Liliane
Xie, Junyu
Sun, Shuyang
Sukthankar, Rahul
Barral, Joëlle K.
Hadsell, Raia
Ghahramani, Zoubin
Zisserman, Andrew
Zhang, Junlin
Sajjadi, Mehdi S. M.
contents Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently solve this task. D4RT utilizes a unified transformer architecture to jointly infer depth, spatio-temporal correspondence, and full camera parameters from a single video. Its core innovation is a novel querying mechanism that sidesteps the heavy computation of dense, per-frame decoding and the complexity of managing multiple, task-specific decoders. Our decoding interface allows the model to independently and flexibly probe the 3D position of any point in space and time. The result is a lightweight and highly scalable method that enables remarkably efficient training and inference. We demonstrate that our approach sets a new state of the art, outperforming previous methods across a wide spectrum of 4D reconstruction tasks. We refer to the project webpage for animated results: https://d4rt-paper.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
Zhang, Chuhan
Moing, Guillaume Le
Koppula, Skanda
Rocco, Ignacio
Momeni, Liliane
Xie, Junyu
Sun, Shuyang
Sukthankar, Rahul
Barral, Joëlle K.
Hadsell, Raia
Ghahramani, Zoubin
Zisserman, Andrew
Zhang, Junlin
Sajjadi, Mehdi S. M.
Computer Vision and Pattern Recognition
Understanding and reconstructing the complex geometry and motion of dynamic scenes from video remains a formidable challenge in computer vision. This paper introduces D4RT, a simple yet powerful feedforward model designed to efficiently solve this task. D4RT utilizes a unified transformer architecture to jointly infer depth, spatio-temporal correspondence, and full camera parameters from a single video. Its core innovation is a novel querying mechanism that sidesteps the heavy computation of dense, per-frame decoding and the complexity of managing multiple, task-specific decoders. Our decoding interface allows the model to independently and flexibly probe the 3D position of any point in space and time. The result is a lightweight and highly scalable method that enables remarkably efficient training and inference. We demonstrate that our approach sets a new state of the art, outperforming previous methods across a wide spectrum of 4D reconstruction tasks. We refer to the project webpage for animated results: https://d4rt-paper.github.io/.
title Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.08924