NeRF-DetS: Enhanced Adaptive Spatial-wise Sampling and View-wise Fusion Strategies for NeRF-based Indoor Multi-view 3D Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912170977001472 |
|---|---|
| author | Huang, Chi Li, Xinyang Qu, Yansong Wu, Changli Li, Xiaofan Zhang, Shengchuan Cao, Liujuan |
| author_facet | Huang, Chi Li, Xinyang Qu, Yansong Wu, Changli Li, Xiaofan Zhang, Shengchuan Cao, Liujuan |
| contents | In indoor scenes, the diverse distribution of object locations and scales makes the visual 3D perception task a big challenge.
Previous works (e.g, NeRF-Det) have demonstrated that implicit representation has the capacity to benefit the visual 3D perception task in indoor scenes with high amount of overlap between input images.
However, previous works cannot fully utilize the advancement of implicit representation because of fixed sampling and simple multi-view feature fusion.
In this paper, inspired by sparse fashion method (e.g, DETR3D), we propose a simple yet effective method, NeRF-DetS, to address above issues. NeRF-DetS includes two modules: Progressive Adaptive Sampling Strategy (PASS) and Depth-Guided Simplified Multi-Head Attention Fusion (DS-MHA).
Specifically,
(1)PASS can automatically sample features of each layer within a dense 3D detector, using offsets predicted by the previous layer.
(2)DS-MHA can not only efficiently fuse multi-view features with strong occlusion awareness but also reduce computational cost.
Extensive experiments on ScanNetV2 dataset demonstrate our NeRF-DetS outperforms NeRF-Det, by achieving +5.02% and +5.92% improvement in mAP under IoU25 and IoU50, respectively. Also, NeRF-DetS shows consistent improvements on ARKITScenes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_13921 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | NeRF-DetS: Enhanced Adaptive Spatial-wise Sampling and View-wise Fusion Strategies for NeRF-based Indoor Multi-view 3D Object Detection Huang, Chi Li, Xinyang Qu, Yansong Wu, Changli Li, Xiaofan Zhang, Shengchuan Cao, Liujuan Computer Vision and Pattern Recognition In indoor scenes, the diverse distribution of object locations and scales makes the visual 3D perception task a big challenge. Previous works (e.g, NeRF-Det) have demonstrated that implicit representation has the capacity to benefit the visual 3D perception task in indoor scenes with high amount of overlap between input images. However, previous works cannot fully utilize the advancement of implicit representation because of fixed sampling and simple multi-view feature fusion. In this paper, inspired by sparse fashion method (e.g, DETR3D), we propose a simple yet effective method, NeRF-DetS, to address above issues. NeRF-DetS includes two modules: Progressive Adaptive Sampling Strategy (PASS) and Depth-Guided Simplified Multi-Head Attention Fusion (DS-MHA). Specifically, (1)PASS can automatically sample features of each layer within a dense 3D detector, using offsets predicted by the previous layer. (2)DS-MHA can not only efficiently fuse multi-view features with strong occlusion awareness but also reduce computational cost. Extensive experiments on ScanNetV2 dataset demonstrate our NeRF-DetS outperforms NeRF-Det, by achieving +5.02% and +5.92% improvement in mAP under IoU25 and IoU50, respectively. Also, NeRF-DetS shows consistent improvements on ARKITScenes. |
| title | NeRF-DetS: Enhanced Adaptive Spatial-wise Sampling and View-wise Fusion Strategies for NeRF-based Indoor Multi-view 3D Object Detection |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2404.13921 |