WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916936162476032 |
|---|---|
| author | Li, Zizun Zhou, Jianjun Wang, Yifan Guo, Haoyu Chang, Wenzheng Zhou, Yang Zhu, Haoyi Chen, Junyi Shen, Chunhua He, Tong |
| author_facet | Li, Zizun Zhou, Jianjun Wang, Yifan Guo, Haoyu Chang, Wenzheng Zhou, Yang Zhu, Haoyi Chen, Junyi Shen, Chunhua He, Tong |
| contents | We present WinT3R, a feed-forward reconstruction model capable of online prediction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To address this, we first introduce a sliding window mechanism that ensures sufficient information exchange among frames within the window, thereby improving the quality of geometric predictions without large computation. In addition, we leverage a compact representation of cameras and maintain a global camera token pool, which enhances the reliability of camera pose estimation without sacrificing efficiency. These designs enable WinT3R to achieve state-of-the-art performance in terms of online reconstruction quality, camera pose estimation, and reconstruction speed, as validated by extensive experiments on diverse datasets. Code and model are publicly available at https://github.com/LiZizun/WinT3R. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_05296 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool Li, Zizun Zhou, Jianjun Wang, Yifan Guo, Haoyu Chang, Wenzheng Zhou, Yang Zhu, Haoyi Chen, Junyi Shen, Chunhua He, Tong Computer Vision and Pattern Recognition Artificial Intelligence We present WinT3R, a feed-forward reconstruction model capable of online prediction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To address this, we first introduce a sliding window mechanism that ensures sufficient information exchange among frames within the window, thereby improving the quality of geometric predictions without large computation. In addition, we leverage a compact representation of cameras and maintain a global camera token pool, which enhances the reliability of camera pose estimation without sacrificing efficiency. These designs enable WinT3R to achieve state-of-the-art performance in terms of online reconstruction quality, camera pose estimation, and reconstruction speed, as validated by extensive experiments on diverse datasets. Code and model are publicly available at https://github.com/LiZizun/WinT3R. |
| title | WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2509.05296 |