Continuous 3D Perception Model with Persistent State

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Qianqian, Zhang, Yifei, Holynski, Aleksander, Efros, Alexei A., Kanazawa, Angjoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917898644094976
author Wang, Qianqian
Zhang, Yifei
Holynski, Aleksander
Efros, Alexei A.
Kanazawa, Angjoo
author_facet Wang, Qianqian
Zhang, Yifei
Holynski, Aleksander
Efros, Alexei A.
Kanazawa, Angjoo
contents We present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this evolving state can be used to generate metric-scale pointmaps (per-pixel 3D points) for each new input in an online fashion. These pointmaps reside within a common coordinate system, and can be accumulated into a coherent, dense scene reconstruction that updates as new images arrive. Our model, called CUT3R (Continuous Updating Transformer for 3D Reconstruction), captures rich priors of real-world scenes: not only can it predict accurate pointmaps from image observations, but it can also infer unseen regions of the scene by probing at virtual, unobserved views. Our method is simple yet highly flexible, naturally accepting varying lengths of images that may be either video streams or unordered photo collections, containing both static and dynamic content. We evaluate our method on various 3D/4D tasks and demonstrate competitive or state-of-the-art performance in each. Project Page: https://cut3r.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2501_12387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Continuous 3D Perception Model with Persistent State
Wang, Qianqian
Zhang, Yifei
Holynski, Aleksander
Efros, Alexei A.
Kanazawa, Angjoo
Computer Vision and Pattern Recognition
We present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this evolving state can be used to generate metric-scale pointmaps (per-pixel 3D points) for each new input in an online fashion. These pointmaps reside within a common coordinate system, and can be accumulated into a coherent, dense scene reconstruction that updates as new images arrive. Our model, called CUT3R (Continuous Updating Transformer for 3D Reconstruction), captures rich priors of real-world scenes: not only can it predict accurate pointmaps from image observations, but it can also infer unseen regions of the scene by probing at virtual, unobserved views. Our method is simple yet highly flexible, naturally accepting varying lengths of images that may be either video streams or unordered photo collections, containing both static and dynamic content. We evaluate our method on various 3D/4D tasks and demonstrate competitive or state-of-the-art performance in each. Project Page: https://cut3r.github.io/
title Continuous 3D Perception Model with Persistent State
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.12387