SpatialTrackerV2: 3D Point Tracking Made Easy
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908456686977024 |
|---|---|
| author | Xiao, Yuxi Wang, Jianyuan Xue, Nan Karaev, Nikita Makarov, Yuri Kang, Bingyi Zhu, Xing Bao, Hujun Shen, Yujun Zhou, Xiaowei |
| author_facet | Xiao, Yuxi Wang, Jianyuan Xue, Nan Karaev, Nikita Makarov, Yuri Kang, Bingyi Zhu, Xing Bao, Hujun Shen, Yujun Zhou, Xiaowei |
| contents | We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50$\times$ faster. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_12462 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SpatialTrackerV2: 3D Point Tracking Made Easy Xiao, Yuxi Wang, Jianyuan Xue, Nan Karaev, Nikita Makarov, Yuri Kang, Bingyi Zhu, Xing Bao, Hujun Shen, Yujun Zhou, Xiaowei Computer Vision and Pattern Recognition We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50$\times$ faster. |
| title | SpatialTrackerV2: 3D Point Tracking Made Easy |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2507.12462 |