Discriminately Treating Motion Components Evolves Joint Depth and Ego-Motion Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Mengtan, Guo, Zizhan, Zhao, Hongbo, Feng, Yi, Xiong, Zuyi, Wang, Yue, Du, Shaoyi, Wang, Hanli, Fan, Rui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917056473989120
author Zhang, Mengtan
Guo, Zizhan
Zhao, Hongbo
Feng, Yi
Xiong, Zuyi
Wang, Yue
Du, Shaoyi
Wang, Hanli
Fan, Rui
author_facet Zhang, Mengtan
Guo, Zizhan
Zhao, Hongbo
Feng, Yi
Xiong, Zuyi
Wang, Yue
Du, Shaoyi
Wang, Hanli
Fan, Rui
contents Unsupervised learning of depth and ego-motion, two fundamental 3D perception tasks, has made significant strides in recent years. However, most methods treat ego-motion as an auxiliary task, either mixing all motion types or excluding depth-independent rotational motions in supervision. Such designs limit the incorporation of strong geometric constraints, reducing reliability and robustness under diverse conditions. This study introduces a discriminative treatment of motion components, leveraging the geometric regularities of their respective rigid flows to benefit both depth and ego-motion estimation. Given consecutive video frames, network outputs first align the optical axes and imaging planes of the source and target cameras. Optical flows between frames are transformed through these alignments, and deviations are quantified to impose geometric constraints individually on each ego-motion component, enabling more targeted refinement. These alignments further reformulate the joint learning process into coaxial and coplanar forms, where depth and each translation component can be mutually derived through closed-form geometric relationships, introducing complementary constraints that improve depth robustness. DiMoDE, a general depth and ego-motion joint learning framework incorporating these designs, achieves state-of-the-art performance on multiple public datasets and a newly collected diverse real-world dataset, particularly under challenging conditions. Our source code will be publicly available at mias.group/DiMoDE upon publication.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discriminately Treating Motion Components Evolves Joint Depth and Ego-Motion Learning
Zhang, Mengtan
Guo, Zizhan
Zhao, Hongbo
Feng, Yi
Xiong, Zuyi
Wang, Yue
Du, Shaoyi
Wang, Hanli
Fan, Rui
Computer Vision and Pattern Recognition
Robotics
Unsupervised learning of depth and ego-motion, two fundamental 3D perception tasks, has made significant strides in recent years. However, most methods treat ego-motion as an auxiliary task, either mixing all motion types or excluding depth-independent rotational motions in supervision. Such designs limit the incorporation of strong geometric constraints, reducing reliability and robustness under diverse conditions. This study introduces a discriminative treatment of motion components, leveraging the geometric regularities of their respective rigid flows to benefit both depth and ego-motion estimation. Given consecutive video frames, network outputs first align the optical axes and imaging planes of the source and target cameras. Optical flows between frames are transformed through these alignments, and deviations are quantified to impose geometric constraints individually on each ego-motion component, enabling more targeted refinement. These alignments further reformulate the joint learning process into coaxial and coplanar forms, where depth and each translation component can be mutually derived through closed-form geometric relationships, introducing complementary constraints that improve depth robustness. DiMoDE, a general depth and ego-motion joint learning framework incorporating these designs, achieves state-of-the-art performance on multiple public datasets and a newly collected diverse real-world dataset, particularly under challenging conditions. Our source code will be publicly available at mias.group/DiMoDE upon publication.
title Discriminately Treating Motion Components Evolves Joint Depth and Ego-Motion Learning
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2511.01502