FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913285551423488 |
|---|---|
| author | Han, Meng Wang, Liang Xiao, Limin Zhang, Hao Zhang, Chenhao Xie, Xilong Zheng, Shuai Dong, Jin |
| author_facet | Han, Meng Wang, Liang Xiao, Limin Zhang, Hao Zhang, Chenhao Xie, Xilong Zheng, Shuai Dong, Jin |
| contents | Point cloud analytics has become a critical workload for embedded and mobile platforms across various applications. Farthest point sampling (FPS) is a fundamental and widely used kernel in point cloud processing. However, the heavy external memory access makes FPS a performance bottleneck for real-time point cloud processing. Although bucket-based farthest point sampling can significantly reduce unnecessary memory accesses during the point sampling stage, the KD-tree construction stage becomes the predominant contributor to execution time. In this paper, we present FuseFPS, an architecture and algorithm co-design for bucket-based farthest point sampling. We first propose a hardware-friendly sampling-driven KD-tree construction algorithm. The algorithm fuses the KD-tree construction stage into the point sampling stage, further reducing memory accesses. Then, we design an efficient accelerator for bucket-based point sampling. The accelerator can offload the entire bucket-based FPS kernel at a low hardware cost. Finally, we evaluate our approach on various point cloud datasets. The detailed experiments show that compared to the state-of-the-art accelerator QuickFPS, FuseFPS achieves about 4.3$\times$ and about 6.1$\times$ improvements on speed and power efficiency, respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_05017 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds Han, Meng Wang, Liang Xiao, Limin Zhang, Hao Zhang, Chenhao Xie, Xilong Zheng, Shuai Dong, Jin Hardware Architecture Point cloud analytics has become a critical workload for embedded and mobile platforms across various applications. Farthest point sampling (FPS) is a fundamental and widely used kernel in point cloud processing. However, the heavy external memory access makes FPS a performance bottleneck for real-time point cloud processing. Although bucket-based farthest point sampling can significantly reduce unnecessary memory accesses during the point sampling stage, the KD-tree construction stage becomes the predominant contributor to execution time. In this paper, we present FuseFPS, an architecture and algorithm co-design for bucket-based farthest point sampling. We first propose a hardware-friendly sampling-driven KD-tree construction algorithm. The algorithm fuses the KD-tree construction stage into the point sampling stage, further reducing memory accesses. Then, we design an efficient accelerator for bucket-based point sampling. The accelerator can offload the entire bucket-based FPS kernel at a low hardware cost. Finally, we evaluate our approach on various point cloud datasets. The detailed experiments show that compared to the state-of-the-art accelerator QuickFPS, FuseFPS achieves about 4.3$\times$ and about 6.1$\times$ improvements on speed and power efficiency, respectively. |
| title | FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds |
| topic | Hardware Architecture |
| url | https://arxiv.org/abs/2309.05017 |