Any4D: Unified Feed-Forward Metric 4D Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karhade, Jay, Keetha, Nikhil, Zhang, Yuchen, Gupta, Tanisha, Sharma, Akash, Scherer, Sebastian, Ramanan, Deva
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914195711197184
author Karhade, Jay
Keetha, Nikhil
Zhang, Yuchen
Gupta, Tanisha
Sharma, Akash
Scherer, Sebastian
Ramanan, Deva
author_facet Karhade, Jay
Keetha, Nikhil
Zhang, Yuchen
Gupta, Tanisha
Sharma, Akash
Scherer, Sebastian
Ramanan, Deva
contents We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N frames, in contrast to prior work that typically focuses on either 2-view dense scene flow or sparse 3D point tracking. Moreover, unlike other recent methods for 4D reconstruction from monocular RGB videos, Any4D can process additional modalities and sensors such as RGB-D frames, IMU-based egomotion, and Radar Doppler measurements, when available. One of the key innovations that allows for such a flexible framework is a modular representation of a 4D scene; specifically, per-view 4D predictions are encoded using a variety of egocentric factors (depthmaps and camera intrinsics) represented in local camera coordinates, and allocentric factors (camera extrinsics and scene flow) represented in global world coordinates. We achieve superior performance across diverse setups - both in terms of accuracy (2-3X lower error) and compute efficiency (15X faster), opening avenues for multiple downstream applications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10935
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Any4D: Unified Feed-Forward Metric 4D Reconstruction
Karhade, Jay
Keetha, Nikhil
Zhang, Yuchen
Gupta, Tanisha
Sharma, Akash
Scherer, Sebastian
Ramanan, Deva
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N frames, in contrast to prior work that typically focuses on either 2-view dense scene flow or sparse 3D point tracking. Moreover, unlike other recent methods for 4D reconstruction from monocular RGB videos, Any4D can process additional modalities and sensors such as RGB-D frames, IMU-based egomotion, and Radar Doppler measurements, when available. One of the key innovations that allows for such a flexible framework is a modular representation of a 4D scene; specifically, per-view 4D predictions are encoded using a variety of egocentric factors (depthmaps and camera intrinsics) represented in local camera coordinates, and allocentric factors (camera extrinsics and scene flow) represented in global world coordinates. We achieve superior performance across diverse setups - both in terms of accuracy (2-3X lower error) and compute efficiency (15X faster), opening avenues for multiple downstream applications.
title Any4D: Unified Feed-Forward Metric 4D Reconstruction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2512.10935