Enhanced Object Tracking by Self-Supervised Auxiliary Depth Estimation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Zhenyu, He, Yujie, Cai, Zhanchuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911884876185600
author Wei, Zhenyu
He, Yujie
Cai, Zhanchuan
author_facet Wei, Zhenyu
He, Yujie
Cai, Zhanchuan
contents RGB-D tracking significantly improves the accuracy of object tracking. However, its dependency on real depth inputs and the complexity involved in multi-modal fusion limit its applicability across various scenarios. The utilization of depth information in RGB-D tracking inspired us to propose a new method, named MDETrack, which trains a tracking network with an additional capability to understand the depth of scenes, through supervised or self-supervised auxiliary Monocular Depth Estimation learning. The outputs of MDETrack's unified feature extractor are fed to the side-by-side tracking head and auxiliary depth estimation head, respectively. The auxiliary module will be discarded in inference, thus keeping the same inference speed. We evaluated our models with various training strategies on multiple datasets, and the results show an improved tracking accuracy even without real depth. Through these findings we highlight the potential of depth estimation in enhancing object tracking performance.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14195
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhanced Object Tracking by Self-Supervised Auxiliary Depth Estimation Learning
Wei, Zhenyu
He, Yujie
Cai, Zhanchuan
Computer Vision and Pattern Recognition
Artificial Intelligence
RGB-D tracking significantly improves the accuracy of object tracking. However, its dependency on real depth inputs and the complexity involved in multi-modal fusion limit its applicability across various scenarios. The utilization of depth information in RGB-D tracking inspired us to propose a new method, named MDETrack, which trains a tracking network with an additional capability to understand the depth of scenes, through supervised or self-supervised auxiliary Monocular Depth Estimation learning. The outputs of MDETrack's unified feature extractor are fed to the side-by-side tracking head and auxiliary depth estimation head, respectively. The auxiliary module will be discarded in inference, thus keeping the same inference speed. We evaluated our models with various training strategies on multiple datasets, and the results show an improved tracking accuracy even without real depth. Through these findings we highlight the potential of depth estimation in enhancing object tracking performance.
title Enhanced Object Tracking by Self-Supervised Auxiliary Depth Estimation Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.14195