SpatialTrackerV2: 3D Point Tracking Made Easy

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiao, Yuxi, Wang, Jianyuan, Xue, Nan, Karaev, Nikita, Makarov, Yuri, Kang, Bingyi, Zhu, Xing, Bao, Hujun, Shen, Yujun, Zhou, Xiaowei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908456686977024
author Xiao, Yuxi
Wang, Jianyuan
Xue, Nan
Karaev, Nikita
Makarov, Yuri
Kang, Bingyi
Zhu, Xing
Bao, Hujun
Shen, Yujun
Zhou, Xiaowei
author_facet Xiao, Yuxi
Wang, Jianyuan
Xue, Nan
Karaev, Nikita
Makarov, Yuri
Kang, Bingyi
Zhu, Xing
Bao, Hujun
Shen, Yujun
Zhou, Xiaowei
contents We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50$\times$ faster.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpatialTrackerV2: 3D Point Tracking Made Easy
Xiao, Yuxi
Wang, Jianyuan
Xue, Nan
Karaev, Nikita
Makarov, Yuri
Kang, Bingyi
Zhu, Xing
Bao, Hujun
Shen, Yujun
Zhou, Xiaowei
Computer Vision and Pattern Recognition
We present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos. Going beyond modular pipelines built on off-the-shelf components for 3D tracking, our approach unifies the intrinsic connections between point tracking, monocular depth, and camera pose estimation into a high-performing and feedforward 3D point tracker. It decomposes world-space 3D motion into scene geometry, camera ego-motion, and pixel-wise object motion, with a fully differentiable and end-to-end architecture, allowing scalable training across a wide range of datasets, including synthetic sequences, posed RGB-D videos, and unlabeled in-the-wild footage. By learning geometry and motion jointly from such heterogeneous data, SpatialTrackerV2 outperforms existing 3D tracking methods by 30%, and matches the accuracy of leading dynamic 3D reconstruction approaches while running 50$\times$ faster.
title SpatialTrackerV2: 3D Point Tracking Made Easy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.12462