2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Cheng-Kun, Chen, Min-Hung, Chuang, Yung-Yu, Lin, Yen-Yu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913201687363584
author Yang, Cheng-Kun
Chen, Min-Hung
Chuang, Yung-Yu
Lin, Yen-Yu
author_facet Yang, Cheng-Kun
Chen, Min-Hung
Chuang, Yung-Yu
Lin, Yen-Yu
contents We present a Multimodal Interlaced Transformer (MIT) that jointly considers 2D and 3D data for weakly supervised point cloud segmentation. Research studies have shown that 2D and 3D features are complementary for point cloud segmentation. However, existing methods require extra 2D annotations to achieve 2D-3D information fusion. Considering the high annotation cost of point clouds, effective 2D and 3D feature fusion based on weakly supervised learning is in great demand. To this end, we propose a transformer model with two encoders and one decoder for weakly supervised point cloud segmentation using only scene-level class tags. Specifically, the two encoders compute the self-attended features for 3D point clouds and 2D multi-view images, respectively. The decoder implements interlaced 2D-3D cross-attention and carries out implicit 2D and 3D feature fusion. We alternately switch the roles of queries and key-value pairs in the decoder layers. It turns out that the 2D and 3D features are iteratively enriched by each other. Experiments show that it performs favorably against existing weakly supervised point cloud segmentation methods by a large margin on the S3DIS and ScanNet benchmarks. The project page will be available at https://jimmy15923.github.io/mit_web/.
format Preprint
id arxiv_https___arxiv_org_abs_2310_12817
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision
Yang, Cheng-Kun
Chen, Min-Hung
Chuang, Yung-Yu
Lin, Yen-Yu
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We present a Multimodal Interlaced Transformer (MIT) that jointly considers 2D and 3D data for weakly supervised point cloud segmentation. Research studies have shown that 2D and 3D features are complementary for point cloud segmentation. However, existing methods require extra 2D annotations to achieve 2D-3D information fusion. Considering the high annotation cost of point clouds, effective 2D and 3D feature fusion based on weakly supervised learning is in great demand. To this end, we propose a transformer model with two encoders and one decoder for weakly supervised point cloud segmentation using only scene-level class tags. Specifically, the two encoders compute the self-attended features for 3D point clouds and 2D multi-view images, respectively. The decoder implements interlaced 2D-3D cross-attention and carries out implicit 2D and 3D feature fusion. We alternately switch the roles of queries and key-value pairs in the decoder layers. It turns out that the 2D and 3D features are iteratively enriched by each other. Experiments show that it performs favorably against existing weakly supervised point cloud segmentation methods by a large margin on the S3DIS and ScanNet benchmarks. The project page will be available at https://jimmy15923.github.io/mit_web/.
title 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level Supervision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2310.12817