Unsupervised Spike Depth Estimation via Cross-modality Cross-domain Knowledge Transfer

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Jiaming, Zhang, Qizhe, Li, Xiaoqi, Li, Jianing, Wang, Guanqun, Lu, Ming, Huang, Tiejun, Zhang, Shanghang
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929427788595200
author Liu, Jiaming
Zhang, Qizhe
Li, Xiaoqi
Li, Jianing
Wang, Guanqun
Lu, Ming
Huang, Tiejun
Zhang, Shanghang
author_facet Liu, Jiaming
Zhang, Qizhe
Li, Xiaoqi
Li, Jianing
Wang, Guanqun
Lu, Ming
Huang, Tiejun
Zhang, Shanghang
contents Neuromorphic spike data, an upcoming modality with high temporal resolution, has shown promising potential in autonomous driving by mitigating the challenges posed by high-velocity motion blur. However, training the spike depth estimation network holds significant challenges in two aspects: sparse spatial information for pixel-wise tasks and difficulties in achieving paired depth labels for temporally intensive spike streams. Therefore, we introduce open-source RGB data to support spike depth estimation, leveraging its annotations and spatial information. The inherent differences in modalities and data distribution make it challenging to directly apply transfer learning from open-source RGB to target spike data. To this end, we propose a cross-modality cross-domain (BiCross) framework to realize unsupervised spike depth estimation by introducing simulated mediate source spike data. Specifically, we design a Coarse-to-Fine Knowledge Distillation (CFKD) approach to facilitate comprehensive cross-modality knowledge transfer while preserving the unique strengths of both modalities, utilizing a spike-oriented uncertainty scheme. Then, we propose a Self-Correcting Teacher-Student (SCTS) mechanism to screen out reliable pixel-wise pseudo labels and ease the domain shift of the student model, which avoids error accumulation in target spike data. To verify the effectiveness of BiCross, we conduct extensive experiments on four scenarios, including Synthetic to Real, Extreme Weather, Scene Changing, and Real Spike. Our method achieves state-of-the-art (SOTA) performances, compared with RGB-oriented unsupervised depth estimation methods. Code and dataset: https://github.com/Theia-4869/BiCross
format Preprint
id arxiv_https___arxiv_org_abs_2208_12527
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Unsupervised Spike Depth Estimation via Cross-modality Cross-domain Knowledge Transfer
Liu, Jiaming
Zhang, Qizhe
Li, Xiaoqi
Li, Jianing
Wang, Guanqun
Lu, Ming
Huang, Tiejun
Zhang, Shanghang
Computer Vision and Pattern Recognition
Neuromorphic spike data, an upcoming modality with high temporal resolution, has shown promising potential in autonomous driving by mitigating the challenges posed by high-velocity motion blur. However, training the spike depth estimation network holds significant challenges in two aspects: sparse spatial information for pixel-wise tasks and difficulties in achieving paired depth labels for temporally intensive spike streams. Therefore, we introduce open-source RGB data to support spike depth estimation, leveraging its annotations and spatial information. The inherent differences in modalities and data distribution make it challenging to directly apply transfer learning from open-source RGB to target spike data. To this end, we propose a cross-modality cross-domain (BiCross) framework to realize unsupervised spike depth estimation by introducing simulated mediate source spike data. Specifically, we design a Coarse-to-Fine Knowledge Distillation (CFKD) approach to facilitate comprehensive cross-modality knowledge transfer while preserving the unique strengths of both modalities, utilizing a spike-oriented uncertainty scheme. Then, we propose a Self-Correcting Teacher-Student (SCTS) mechanism to screen out reliable pixel-wise pseudo labels and ease the domain shift of the student model, which avoids error accumulation in target spike data. To verify the effectiveness of BiCross, we conduct extensive experiments on four scenarios, including Synthetic to Real, Extreme Weather, Scene Changing, and Real Spike. Our method achieves state-of-the-art (SOTA) performances, compared with RGB-oriented unsupervised depth estimation methods. Code and dataset: https://github.com/Theia-4869/BiCross
title Unsupervised Spike Depth Estimation via Cross-modality Cross-domain Knowledge Transfer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2208.12527