STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lan, Yushi, Luo, Yihang, Hong, Fangzhou, Zhou, Shangchen, Chen, Honghua, Lyu, Zhaoyang, Yang, Shuai, Dai, Bo, Loy, Chen Change, Pan, Xingang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909737055944704
author Lan, Yushi
Luo, Yihang
Hong, Fangzhou
Zhou, Shangchen
Chen, Honghua
Lyu, Zhaoyang
Yang, Shuai
Dai, Bo
Loy, Chen Change
Pan, Xingang
author_facet Lan, Yushi
Luo, Yihang
Hong, Fangzhou
Zhou, Shangchen
Chen, Honghua
Lyu, Zhaoyang
Yang, Shuai
Dai, Bo
Loy, Chen Change
Pan, Xingang
contents We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-view reconstruction either depend on expensive global optimization or rely on simplistic memory mechanisms that scale poorly with sequence length. In contrast, STream3R introduces an streaming framework that processes image sequences efficiently using causal attention, inspired by advances in modern language modeling. By learning geometric priors from large-scale 3D datasets, STream3R generalizes well to diverse and challenging scenarios, including dynamic scenes where traditional methods often fail. Extensive experiments show that our method consistently outperforms prior work across both static and dynamic scene benchmarks. Moreover, STream3R is inherently compatible with LLM-style training infrastructure, enabling efficient large-scale pretraining and fine-tuning for various downstream 3D tasks. Our results underscore the potential of causal Transformer models for online 3D perception, paving the way for real-time 3D understanding in streaming environments. More details can be found in our project page: https://nirvanalan.github.io/projects/stream3r.
format Preprint
id arxiv_https___arxiv_org_abs_2508_10893
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
Lan, Yushi
Luo, Yihang
Hong, Fangzhou
Zhou, Shangchen
Chen, Honghua
Lyu, Zhaoyang
Yang, Shuai
Dai, Bo
Loy, Chen Change
Pan, Xingang
Computer Vision and Pattern Recognition
We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-view reconstruction either depend on expensive global optimization or rely on simplistic memory mechanisms that scale poorly with sequence length. In contrast, STream3R introduces an streaming framework that processes image sequences efficiently using causal attention, inspired by advances in modern language modeling. By learning geometric priors from large-scale 3D datasets, STream3R generalizes well to diverse and challenging scenarios, including dynamic scenes where traditional methods often fail. Extensive experiments show that our method consistently outperforms prior work across both static and dynamic scene benchmarks. Moreover, STream3R is inherently compatible with LLM-style training infrastructure, enabling efficient large-scale pretraining and fine-tuning for various downstream 3D tasks. Our results underscore the potential of causal Transformer models for online 3D perception, paving the way for real-time 3D understanding in streaming environments. More details can be found in our project page: https://nirvanalan.github.io/projects/stream3r.
title STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.10893