Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Tao, Yang, Peishan, Jin, Yudong, Cai, Yingfeng, Yin, Wei, Ren, Weiqiang, Zhang, Qian, Hua, Wei, Peng, Sida, Guo, Xiaoyang, Zhou, Xiaowei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915927845502976
author Xie, Tao
Yang, Peishan
Jin, Yudong
Cai, Yingfeng
Yin, Wei
Ren, Weiqiang
Zhang, Qian
Hua, Wei
Peng, Sida
Guo, Xiaoyang
Zhou, Xiaowei
author_facet Xie, Tao
Yang, Peishan
Jin, Yudong
Cai, Yingfeng
Yin, Wei
Ren, Weiqiang
Zhang, Qian
Hua, Wei
Peng, Sida
Guo, Xiaoyang
Zhou, Xiaowei
contents This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D priors or geometric constraints. However, these methods often struggle to maintain reconstruction accuracy and consistency over long sequences due to limited memory capacity and the inability to effectively capture global contextual cues. In contrast, humans can naturally exploit the global understanding of the scene to inform local perception. Motivated by this, we propose a novel neural global context representation that efficiently compresses and retains long-range scene information, enabling the model to leverage extensive contextual cues for enhanced reconstruction accuracy and consistency. The context representation is realized through a set of lightweight neural sub-networks that are rapidly adapted during test time via self-supervised objectives, which substantially increases memory capacity without incurring significant computational overhead. The experiments on multiple large-scale benchmarks, including the KITTI Odometry~\cite{Geiger2012CVPR} and Oxford Spires~\cite{tao2025spires} datasets, demonstrate the effectiveness of our approach in handling ultra-large scenes, achieving leading pose accuracy and state-of-the-art 3D reconstruction accuracy while maintaining efficiency. Code is available at https://zju3dv.github.io/scal3r.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08542
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
Xie, Tao
Yang, Peishan
Jin, Yudong
Cai, Yingfeng
Yin, Wei
Ren, Weiqiang
Zhang, Qian
Hua, Wei
Peng, Sida
Guo, Xiaoyang
Zhou, Xiaowei
Computer Vision and Pattern Recognition
This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D priors or geometric constraints. However, these methods often struggle to maintain reconstruction accuracy and consistency over long sequences due to limited memory capacity and the inability to effectively capture global contextual cues. In contrast, humans can naturally exploit the global understanding of the scene to inform local perception. Motivated by this, we propose a novel neural global context representation that efficiently compresses and retains long-range scene information, enabling the model to leverage extensive contextual cues for enhanced reconstruction accuracy and consistency. The context representation is realized through a set of lightweight neural sub-networks that are rapidly adapted during test time via self-supervised objectives, which substantially increases memory capacity without incurring significant computational overhead. The experiments on multiple large-scale benchmarks, including the KITTI Odometry~\cite{Geiger2012CVPR} and Oxford Spires~\cite{tao2025spires} datasets, demonstrate the effectiveness of our approach in handling ultra-large scenes, achieving leading pose accuracy and state-of-the-art 3D reconstruction accuracy while maintaining efficiency. Code is available at https://zju3dv.github.io/scal3r.
title Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.08542