$SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yinshuang, Chen, Dian, Liu, Katherine, Zakharov, Sergey, Ambrus, Rares, Daniilidis, Kostas, Guizilini, Vitor
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909385762013184
author Xu, Yinshuang
Chen, Dian
Liu, Katherine
Zakharov, Sergey
Ambrus, Rares
Daniilidis, Kostas
Guizilini, Vitor
author_facet Xu, Yinshuang
Chen, Dian
Liu, Katherine
Zakharov, Sergey
Ambrus, Rares
Daniilidis, Kostas
Guizilini, Vitor
contents Incorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D learning. Equivariance serves as a valuable inductive prior, aiding in the generation of robust multi-view features for 3D scene understanding. In this paper, we explore the application of equivariant multi-view learning to depth estimation, not only recognizing its significance for computer vision and robotics but also addressing the limitations of previous research. Most prior studies have either overlooked equivariance in this setting or achieved only approximate equivariance through data augmentation, which often leads to inconsistencies across different reference frames. To address this issue, we propose to embed $SE(3)$ equivariance into the Perceiver IO architecture. We employ Spherical Harmonics for positional encoding to ensure 3D rotation equivariance, and develop a specialized equivariant encoder and decoder within the Perceiver IO architecture. To validate our model, we applied it to the task of stereo depth estimation, achieving state of the art results on real-world datasets without explicit geometric constraints or extensive data augmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07326
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle $SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation
Xu, Yinshuang
Chen, Dian
Liu, Katherine
Zakharov, Sergey
Ambrus, Rares
Daniilidis, Kostas
Guizilini, Vitor
Computer Vision and Pattern Recognition
Incorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D learning. Equivariance serves as a valuable inductive prior, aiding in the generation of robust multi-view features for 3D scene understanding. In this paper, we explore the application of equivariant multi-view learning to depth estimation, not only recognizing its significance for computer vision and robotics but also addressing the limitations of previous research. Most prior studies have either overlooked equivariance in this setting or achieved only approximate equivariance through data augmentation, which often leads to inconsistencies across different reference frames. To address this issue, we propose to embed $SE(3)$ equivariance into the Perceiver IO architecture. We employ Spherical Harmonics for positional encoding to ensure 3D rotation equivariance, and develop a specialized equivariant encoder and decoder within the Perceiver IO architecture. To validate our model, we applied it to the task of stereo depth estimation, achieving state of the art results on real-world datasets without explicit geometric constraints or extensive data augmentation.
title $SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.07326