Zero-Shot Monocular Scene Flow Estimation in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Yiqing, Badki, Abhishek, Su, Hang, Tompkin, James, Gallo, Orazio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913655454433280
author Liang, Yiqing
Badki, Abhishek
Su, Hang
Tompkin, James
Gallo, Orazio
author_facet Liang, Yiqing
Badki, Abhishek
Su, Hang
Tompkin, James
Gallo, Orazio
contents Large models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow. Even though scene flow has wide potential use, it is not used in practice because current predictive models do not generalize well. We identify three key challenges and propose solutions for each. First, we create a method that jointly estimates geometry and motion for accurate prediction. Second, we alleviate scene flow data scarcity with a data recipe that affords us 1M annotated training samples across diverse synthetic scenes. Third, we evaluate different parameterizations for scene flow prediction and adopt a natural and effective parameterization. Our resulting model outperforms existing methods as well as baselines built on large-scale models in terms of 3D end-point error, and shows zero-shot generalization to the casually captured videos from DAVIS and the robotic manipulation scenes from RoboTAP. Overall, our approach makes scene flow prediction more practical in-the-wild.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10357
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Monocular Scene Flow Estimation in the Wild
Liang, Yiqing
Badki, Abhishek
Su, Hang
Tompkin, James
Gallo, Orazio
Computer Vision and Pattern Recognition
Large models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow. Even though scene flow has wide potential use, it is not used in practice because current predictive models do not generalize well. We identify three key challenges and propose solutions for each. First, we create a method that jointly estimates geometry and motion for accurate prediction. Second, we alleviate scene flow data scarcity with a data recipe that affords us 1M annotated training samples across diverse synthetic scenes. Third, we evaluate different parameterizations for scene flow prediction and adopt a natural and effective parameterization. Our resulting model outperforms existing methods as well as baselines built on large-scale models in terms of 3D end-point error, and shows zero-shot generalization to the casually captured videos from DAVIS and the robotic manipulation scenes from RoboTAP. Overall, our approach makes scene flow prediction more practical in-the-wild.
title Zero-Shot Monocular Scene Flow Estimation in the Wild
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.10357