Seurat: From Moving Points to Depth

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cho, Seokju, Huang, Jiahui, Kim, Seungryong, Lee, Joon-Young
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913801102688256
author Cho, Seokju
Huang, Jiahui
Kim, Seungryong
Lee, Joon-Young
author_facet Cho, Seokju
Huang, Jiahui
Kim, Seungryong
Lee, Joon-Young
contents Accurate depth estimation from monocular videos remains challenging due to ambiguities inherent in single-view geometry, as crucial depth cues like stereopsis are absent. However, humans often perceive relative depth intuitively by observing variations in the size and spacing of objects as they move. Inspired by this, we propose a novel method that infers relative depth by examining the spatial relationships and temporal evolution of a set of tracked 2D trajectories. Specifically, we use off-the-shelf point tracking models to capture 2D trajectories. Then, our approach employs spatial and temporal transformers to process these trajectories and directly infer depth changes over time. Evaluated on the TAPVid-3D benchmark, our method demonstrates robust zero-shot performance, generalizing effectively from synthetic to real-world datasets. Results indicate that our approach achieves temporally smooth, high-accuracy depth predictions across diverse domains.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seurat: From Moving Points to Depth
Cho, Seokju
Huang, Jiahui
Kim, Seungryong
Lee, Joon-Young
Computer Vision and Pattern Recognition
Accurate depth estimation from monocular videos remains challenging due to ambiguities inherent in single-view geometry, as crucial depth cues like stereopsis are absent. However, humans often perceive relative depth intuitively by observing variations in the size and spacing of objects as they move. Inspired by this, we propose a novel method that infers relative depth by examining the spatial relationships and temporal evolution of a set of tracked 2D trajectories. Specifically, we use off-the-shelf point tracking models to capture 2D trajectories. Then, our approach employs spatial and temporal transformers to process these trajectories and directly infer depth changes over time. Evaluated on the TAPVid-3D benchmark, our method demonstrates robust zero-shot performance, generalizing effectively from synthetic to real-world datasets. Results indicate that our approach achieves temporally smooth, high-accuracy depth predictions across diverse domains.
title Seurat: From Moving Points to Depth
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.14687