Deep Cost Ray Fusion for Sparse Depth Video Completion

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Jungeon, Kim, Soongjin, Park, Jaesik, Lee, Seungyong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910617453985792
author Kim, Jungeon
Kim, Soongjin
Park, Jaesik
Lee, Seungyong
author_facet Kim, Jungeon
Kim, Soongjin
Park, Jaesik
Lee, Seungyong
contents In this paper, we present a learning-based framework for sparse depth video completion. Given a sparse depth map and a color image at a certain viewpoint, our approach makes a cost volume that is constructed on depth hypothesis planes. To effectively fuse sequential cost volumes of the multiple viewpoints for improved depth completion, we introduce a learning-based cost volume fusion framework, namely RayFusion, that effectively leverages the attention mechanism for each pair of overlapped rays in adjacent cost volumes. As a result of leveraging feature statistics accumulated over time, our proposed framework consistently outperforms or rivals state-of-the-art approaches on diverse indoor and outdoor datasets, including the KITTI Depth Completion benchmark, VOID Depth Completion benchmark, and ScanNetV2 dataset, using much fewer network parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2409_14935
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deep Cost Ray Fusion for Sparse Depth Video Completion
Kim, Jungeon
Kim, Soongjin
Park, Jaesik
Lee, Seungyong
Computer Vision and Pattern Recognition
In this paper, we present a learning-based framework for sparse depth video completion. Given a sparse depth map and a color image at a certain viewpoint, our approach makes a cost volume that is constructed on depth hypothesis planes. To effectively fuse sequential cost volumes of the multiple viewpoints for improved depth completion, we introduce a learning-based cost volume fusion framework, namely RayFusion, that effectively leverages the attention mechanism for each pair of overlapped rays in adjacent cost volumes. As a result of leveraging feature statistics accumulated over time, our proposed framework consistently outperforms or rivals state-of-the-art approaches on diverse indoor and outdoor datasets, including the KITTI Depth Completion benchmark, VOID Depth Completion benchmark, and ScanNetV2 dataset, using much fewer network parameters.
title Deep Cost Ray Fusion for Sparse Depth Video Completion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.14935