MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913575091568640 |
|---|---|
| author | Tran, Duc Dang Trung Kang, Byeongkeun Lee, Yeejin |
| author_facet | Tran, Duc Dang Trung Kang, Byeongkeun Lee, Yeejin |
| contents | Recently, transformer-based techniques incorporating superpoints have become prevalent in 3D instance segmentation. However, they often encounter an over-segmentation problem, especially noticeable with large objects. Additionally, unreliable mask predictions stemming from superpoint mask prediction further compound this issue. To address these challenges, we propose a novel framework called MSTA3D. It leverages multi-scale feature representation and introduces a twin-attention mechanism to effectively capture them. Furthermore, MSTA3D integrates a box query with a box regularizer, offering a complementary spatial constraint alongside semantic queries. Experimental evaluations on ScanNetV2, ScanNet200 and S3DIS datasets demonstrate that our approach surpasses state-of-the-art 3D instance segmentation methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_01781 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation Tran, Duc Dang Trung Kang, Byeongkeun Lee, Yeejin Computer Vision and Pattern Recognition I.2.10 Recently, transformer-based techniques incorporating superpoints have become prevalent in 3D instance segmentation. However, they often encounter an over-segmentation problem, especially noticeable with large objects. Additionally, unreliable mask predictions stemming from superpoint mask prediction further compound this issue. To address these challenges, we propose a novel framework called MSTA3D. It leverages multi-scale feature representation and introduces a twin-attention mechanism to effectively capture them. Furthermore, MSTA3D integrates a box query with a box regularizer, offering a complementary spatial constraint alongside semantic queries. Experimental evaluations on ScanNetV2, ScanNet200 and S3DIS datasets demonstrate that our approach surpasses state-of-the-art 3D instance segmentation methods. |
| title | MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation |
| topic | Computer Vision and Pattern Recognition I.2.10 |
| url | https://arxiv.org/abs/2411.01781 |