VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Kaichen, Chen, Changhao, Wang, Bing, Saputra, Muhamad Risqi U., Trigoni, Niki, Markham, Andrew
Format: Preprint
Publié: 2020
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910788556423168
author Zhou, Kaichen
Chen, Changhao
Wang, Bing
Saputra, Muhamad Risqi U.
Trigoni, Niki
Markham, Andrew
author_facet Zhou, Kaichen
Chen, Changhao
Wang, Bing
Saputra, Muhamad Risqi U.
Trigoni, Niki
Markham, Andrew
contents Recent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less well studied. In particular, we note that previous approaches towards deep fusion do not perform significantly better than models employing a single modality. We conjecture that this is because of the naive approaches to feature space fusion through summation or concatenation which do not take into account the different strengths of each modality. To address this, we propose an end-to-end framework, termed VMLoc, to fuse different sensor inputs into a common latent space through a variational Product-of-Experts (PoE) followed by attention-based fusion. Unlike previous multimodal variational works directly adapting the objective function of vanilla variational auto-encoder, we show how camera localization can be accurately estimated through an unbiased objective function based on importance weighting. Our model is extensively evaluated on RGB-D datasets and the results prove the efficacy of our model. The source code is available at https://github.com/kaichen-z/VMLoc.
format Preprint
id arxiv_https___arxiv_org_abs_2003_07289
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
Zhou, Kaichen
Chen, Changhao
Wang, Bing
Saputra, Muhamad Risqi U.
Trigoni, Niki
Markham, Andrew
Computer Vision and Pattern Recognition
Image and Video Processing
Recent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less well studied. In particular, we note that previous approaches towards deep fusion do not perform significantly better than models employing a single modality. We conjecture that this is because of the naive approaches to feature space fusion through summation or concatenation which do not take into account the different strengths of each modality. To address this, we propose an end-to-end framework, termed VMLoc, to fuse different sensor inputs into a common latent space through a variational Product-of-Experts (PoE) followed by attention-based fusion. Unlike previous multimodal variational works directly adapting the objective function of vanilla variational auto-encoder, we show how camera localization can be accurately estimated through an unbiased objective function based on importance weighting. Our model is extensively evaluated on RGB-D datasets and the results prove the efficacy of our model. The source code is available at https://github.com/kaichen-z/VMLoc.
title VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2003.07289