Slide FFT on a homogeneous mesh in wafer-scale computing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913191783563264 |
|---|---|
| author | van Putten, Maurice H. P. M. Wilson, Leighton Lavely, Adam W. Hair, Mark |
| author_facet | van Putten, Maurice H. P. M. Wilson, Leighton Lavely, Adam W. Hair, Mark |
| contents | Searches for signals at low signal-to-noise ratios frequently involve the Fast Fourier Transform (FFT). For high-throughput searches, we here consider FFT on the homogeneous mesh of Processing Elements (PEs) of a wafer-scale engine (WSE). To minimize memory overhead in the inherently non-local FFT algorithm, we introduce a new synchronous slide operation ({\em Slide}) exploiting the fast interconnect between adjacent PEs. Feasibility of compute-limited performance is demonstrated in linear scaling of Slide execution times with varying array size in preliminary benchmarks on the CS-2 WSE. The proposed implementation appears opportune to accelerate and open the full discovery potential of FFT-based signal processing in multi-messenger astronomy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_05427 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Slide FFT on a homogeneous mesh in wafer-scale computing van Putten, Maurice H. P. M. Wilson, Leighton Lavely, Adam W. Hair, Mark Signal Processing Instrumentation and Methods for Astrophysics Distributed, Parallel, and Cluster Computing Data Structures and Algorithms Searches for signals at low signal-to-noise ratios frequently involve the Fast Fourier Transform (FFT). For high-throughput searches, we here consider FFT on the homogeneous mesh of Processing Elements (PEs) of a wafer-scale engine (WSE). To minimize memory overhead in the inherently non-local FFT algorithm, we introduce a new synchronous slide operation ({\em Slide}) exploiting the fast interconnect between adjacent PEs. Feasibility of compute-limited performance is demonstrated in linear scaling of Slide execution times with varying array size in preliminary benchmarks on the CS-2 WSE. The proposed implementation appears opportune to accelerate and open the full discovery potential of FFT-based signal processing in multi-messenger astronomy. |
| title | Slide FFT on a homogeneous mesh in wafer-scale computing |
| topic | Signal Processing Instrumentation and Methods for Astrophysics Distributed, Parallel, and Cluster Computing Data Structures and Algorithms |
| url | https://arxiv.org/abs/2401.05427 |