Dissecting RGB-D Learning for Improved Multi-modal Fusion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Hao, Zhou, Haoran, Zhang, Yunshu, Lin, Zheng, Deng, Yongjian
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908408128471040
author Chen, Hao
Zhou, Haoran
Zhang, Yunshu
Lin, Zheng
Deng, Yongjian
author_facet Chen, Hao
Zhou, Haoran
Zhang, Yunshu
Lin, Zheng
Deng, Yongjian
contents In the RGB-D vision community, extensive research has been focused on designing multi-modal learning strategies and fusion structures. However, the complementary and fusion mechanisms in RGB-D models remain a black box. In this paper, we present an analytical framework and a novel score to dissect the RGB-D vision community. Our approach involves measuring proposed semantic variance and feature similarity across modalities and levels, conducting visual and quantitative analyzes on multi-modal learning through comprehensive experiments. Specifically, we investigate the consistency and specialty of features across modalities, evolution rules within each modality, and the collaboration logic used when optimizing a RGB-D model. Our studies reveal/verify several important findings, such as the discrepancy in cross-modal features and the hybrid multi-modal cooperation rule, which highlights consistency and specialty simultaneously for complementary inference. We also showcase the versatility of the proposed RGB-D dissection method and introduce a straightforward fusion strategy based on our findings, which delivers significant enhancements across various tasks and even other multi-modal data.
format Preprint
id arxiv_https___arxiv_org_abs_2308_10019
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dissecting RGB-D Learning for Improved Multi-modal Fusion
Chen, Hao
Zhou, Haoran
Zhang, Yunshu
Lin, Zheng
Deng, Yongjian
Computer Vision and Pattern Recognition
In the RGB-D vision community, extensive research has been focused on designing multi-modal learning strategies and fusion structures. However, the complementary and fusion mechanisms in RGB-D models remain a black box. In this paper, we present an analytical framework and a novel score to dissect the RGB-D vision community. Our approach involves measuring proposed semantic variance and feature similarity across modalities and levels, conducting visual and quantitative analyzes on multi-modal learning through comprehensive experiments. Specifically, we investigate the consistency and specialty of features across modalities, evolution rules within each modality, and the collaboration logic used when optimizing a RGB-D model. Our studies reveal/verify several important findings, such as the discrepancy in cross-modal features and the hybrid multi-modal cooperation rule, which highlights consistency and specialty simultaneously for complementary inference. We also showcase the versatility of the proposed RGB-D dissection method and introduce a straightforward fusion strategy based on our findings, which delivers significant enhancements across various tasks and even other multi-modal data.
title Dissecting RGB-D Learning for Improved Multi-modal Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.10019