WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Zizun, Zhou, Jianjun, Wang, Yifan, Guo, Haoyu, Chang, Wenzheng, Zhou, Yang, Zhu, Haoyi, Chen, Junyi, Shen, Chunhua, He, Tong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916936162476032
author Li, Zizun
Zhou, Jianjun
Wang, Yifan
Guo, Haoyu
Chang, Wenzheng
Zhou, Yang
Zhu, Haoyi
Chen, Junyi
Shen, Chunhua
He, Tong
author_facet Li, Zizun
Zhou, Jianjun
Wang, Yifan
Guo, Haoyu
Chang, Wenzheng
Zhou, Yang
Zhu, Haoyi
Chen, Junyi
Shen, Chunhua
He, Tong
contents We present WinT3R, a feed-forward reconstruction model capable of online prediction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To address this, we first introduce a sliding window mechanism that ensures sufficient information exchange among frames within the window, thereby improving the quality of geometric predictions without large computation. In addition, we leverage a compact representation of cameras and maintain a global camera token pool, which enhances the reliability of camera pose estimation without sacrificing efficiency. These designs enable WinT3R to achieve state-of-the-art performance in terms of online reconstruction quality, camera pose estimation, and reconstruction speed, as validated by extensive experiments on diverse datasets. Code and model are publicly available at https://github.com/LiZizun/WinT3R.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05296
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
Li, Zizun
Zhou, Jianjun
Wang, Yifan
Guo, Haoyu
Chang, Wenzheng
Zhou, Yang
Zhu, Haoyi
Chen, Junyi
Shen, Chunhua
He, Tong
Computer Vision and Pattern Recognition
Artificial Intelligence
We present WinT3R, a feed-forward reconstruction model capable of online prediction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To address this, we first introduce a sliding window mechanism that ensures sufficient information exchange among frames within the window, thereby improving the quality of geometric predictions without large computation. In addition, we leverage a compact representation of cameras and maintain a global camera token pool, which enhances the reliability of camera pose estimation without sacrificing efficiency. These designs enable WinT3R to achieve state-of-the-art performance in terms of online reconstruction quality, camera pose estimation, and reconstruction speed, as validated by extensive experiments on diverse datasets. Code and model are publicly available at https://github.com/LiZizun/WinT3R.
title WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.05296