MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tran, Duc Dang Trung, Kang, Byeongkeun, Lee, Yeejin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913575091568640
author Tran, Duc Dang Trung
Kang, Byeongkeun
Lee, Yeejin
author_facet Tran, Duc Dang Trung
Kang, Byeongkeun
Lee, Yeejin
contents Recently, transformer-based techniques incorporating superpoints have become prevalent in 3D instance segmentation. However, they often encounter an over-segmentation problem, especially noticeable with large objects. Additionally, unreliable mask predictions stemming from superpoint mask prediction further compound this issue. To address these challenges, we propose a novel framework called MSTA3D. It leverages multi-scale feature representation and introduces a twin-attention mechanism to effectively capture them. Furthermore, MSTA3D integrates a box query with a box regularizer, offering a complementary spatial constraint alongside semantic queries. Experimental evaluations on ScanNetV2, ScanNet200 and S3DIS datasets demonstrate that our approach surpasses state-of-the-art 3D instance segmentation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01781
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
Tran, Duc Dang Trung
Kang, Byeongkeun
Lee, Yeejin
Computer Vision and Pattern Recognition
I.2.10
Recently, transformer-based techniques incorporating superpoints have become prevalent in 3D instance segmentation. However, they often encounter an over-segmentation problem, especially noticeable with large objects. Additionally, unreliable mask predictions stemming from superpoint mask prediction further compound this issue. To address these challenges, we propose a novel framework called MSTA3D. It leverages multi-scale feature representation and introduces a twin-attention mechanism to effectively capture them. Furthermore, MSTA3D integrates a box query with a box regularizer, offering a complementary spatial constraint alongside semantic queries. Experimental evaluations on ScanNetV2, ScanNet200 and S3DIS datasets demonstrate that our approach surpasses state-of-the-art 3D instance segmentation methods.
title MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
topic Computer Vision and Pattern Recognition
I.2.10
url https://arxiv.org/abs/2411.01781