Project-and-Fuse: Improving RGB-D Semantic Segmentation via Graph Convolution Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jiang, Xiaoyan, Wang, Bohan, Wan, Xinlong, Chen, Shanshan, Fujita, Hamido, Juaid, Hanan Abd. Al
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910925603209216
author Jiang, Xiaoyan
Wang, Bohan
Wan, Xinlong
Chen, Shanshan
Fujita, Hamido
Juaid, Hanan Abd. Al
author_facet Jiang, Xiaoyan
Wang, Bohan
Wan, Xinlong
Chen, Shanshan
Fujita, Hamido
Juaid, Hanan Abd. Al
contents Most existing RGB-D semantic segmentation methods focus on the feature level fusion, including complex cross-modality and cross-scale fusion modules. However, these methods may cause misalignment problem in the feature fusion process and counter-intuitive patches in the segmentation results. Inspired by the popular pixel-node-pixel pipeline, we propose to 1) fuse features from two modalities in a late fusion style, during which the geometric feature injection is guided by texture feature prior; 2) employ Graph Neural Networks (GNNs) on the fused feature to alleviate the emergence of irregular patches by inferring patch relationship. At the 3D feature extraction stage, we argue that traditional CNNs are not efficient enough for depth maps. So, we encode depth map into normal map, after which CNNs can easily extract object surface tendencies.At projection matrix generation stage, we find the existence of Biased-Assignment and Ambiguous-Locality issues in the original pipeline. Therefore, we propose to 1) adopt the Kullback-Leibler Loss to ensure no missing important pixel features, which can be viewed as hard pixel mining process; 2) connect regions that are close to each other in the Euclidean space as well as in the semantic space with larger edge weights so that location informations can been considered. Extensive experiments on two public datasets, NYU-DepthV2 and SUN RGB-D, have shown that our approach can consistently boost the performance of RGB-D semantic segmentation task.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Project-and-Fuse: Improving RGB-D Semantic Segmentation via Graph Convolution Networks
Jiang, Xiaoyan
Wang, Bohan
Wan, Xinlong
Chen, Shanshan
Fujita, Hamido
Juaid, Hanan Abd. Al
Computer Vision and Pattern Recognition
Most existing RGB-D semantic segmentation methods focus on the feature level fusion, including complex cross-modality and cross-scale fusion modules. However, these methods may cause misalignment problem in the feature fusion process and counter-intuitive patches in the segmentation results. Inspired by the popular pixel-node-pixel pipeline, we propose to 1) fuse features from two modalities in a late fusion style, during which the geometric feature injection is guided by texture feature prior; 2) employ Graph Neural Networks (GNNs) on the fused feature to alleviate the emergence of irregular patches by inferring patch relationship. At the 3D feature extraction stage, we argue that traditional CNNs are not efficient enough for depth maps. So, we encode depth map into normal map, after which CNNs can easily extract object surface tendencies.At projection matrix generation stage, we find the existence of Biased-Assignment and Ambiguous-Locality issues in the original pipeline. Therefore, we propose to 1) adopt the Kullback-Leibler Loss to ensure no missing important pixel features, which can be viewed as hard pixel mining process; 2) connect regions that are close to each other in the Euclidean space as well as in the semantic space with larger edge weights so that location informations can been considered. Extensive experiments on two public datasets, NYU-DepthV2 and SUN RGB-D, have shown that our approach can consistently boost the performance of RGB-D semantic segmentation task.
title Project-and-Fuse: Improving RGB-D Semantic Segmentation via Graph Convolution Networks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.18851