AugUndo: Scaling Up Augmentations for Monocular Depth Completion and Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yangchao, Liu, Tian Yu, Park, Hyoungseob, Soatto, Stefano, Lao, Dong, Wong, Alex
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929428136722432
author Wu, Yangchao
Liu, Tian Yu
Park, Hyoungseob
Soatto, Stefano
Lao, Dong
Wong, Alex
author_facet Wu, Yangchao
Liu, Tian Yu
Park, Hyoungseob
Soatto, Stefano
Lao, Dong
Wong, Alex
contents Unsupervised depth completion and estimation methods are trained by minimizing reconstruction error. Block artifacts from resampling, intensity saturation, and occlusions are amongst the many undesirable by-products of common data augmentation schemes that affect image reconstruction quality, and thus the training signal. Hence, typical augmentations on images viewed as essential to training pipelines in other vision tasks have seen limited use beyond small image intensity changes and flipping. The sparse depth modality in depth completion have seen even less use as intensity transformations alter the scale of the 3D scene, and geometric transformations may decimate the sparse points during resampling. We propose a method that unlocks a wide range of previously-infeasible geometric augmentations for unsupervised depth completion and estimation. This is achieved by reversing, or ``undo''-ing, geometric transformations to the coordinates of the output depth, warping the depth map back to the original reference frame. This enables computing the reconstruction losses using the original images and sparse depth maps, eliminating the pitfalls of naive loss computation on the augmented inputs and allowing us to scale up augmentations to boost performance. We demonstrate our method on indoor (VOID) and outdoor (KITTI) datasets, where we consistently improve upon recent methods across both datasets as well as generalization to four other datasets. Code available at: https://github.com/alexklwong/augundo.
format Preprint
id arxiv_https___arxiv_org_abs_2310_09739
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AugUndo: Scaling Up Augmentations for Monocular Depth Completion and Estimation
Wu, Yangchao
Liu, Tian Yu
Park, Hyoungseob
Soatto, Stefano
Lao, Dong
Wong, Alex
Computer Vision and Pattern Recognition
Unsupervised depth completion and estimation methods are trained by minimizing reconstruction error. Block artifacts from resampling, intensity saturation, and occlusions are amongst the many undesirable by-products of common data augmentation schemes that affect image reconstruction quality, and thus the training signal. Hence, typical augmentations on images viewed as essential to training pipelines in other vision tasks have seen limited use beyond small image intensity changes and flipping. The sparse depth modality in depth completion have seen even less use as intensity transformations alter the scale of the 3D scene, and geometric transformations may decimate the sparse points during resampling. We propose a method that unlocks a wide range of previously-infeasible geometric augmentations for unsupervised depth completion and estimation. This is achieved by reversing, or ``undo''-ing, geometric transformations to the coordinates of the output depth, warping the depth map back to the original reference frame. This enables computing the reconstruction losses using the original images and sparse depth maps, eliminating the pitfalls of naive loss computation on the augmented inputs and allowing us to scale up augmentations to boost performance. We demonstrate our method on indoor (VOID) and outdoor (KITTI) datasets, where we consistently improve upon recent methods across both datasets as well as generalization to four other datasets. Code available at: https://github.com/alexklwong/augundo.
title AugUndo: Scaling Up Augmentations for Monocular Depth Completion and Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.09739