A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917240280973312 |
|---|---|
| author | Sapkas, M. Triossi, A. Zanetti, M. |
| author_facet | Sapkas, M. Triossi, A. Zanetti, M. |
| contents | This work explores the use of the AMD Xilinx Versal Adaptable Intelligent Engine (AIE) to accelerate Gated Recurrent Unit (GRU) inference for latency constrained applications. We present a custom workload distribution framework across the AIE's vector processors and propose a hybrid AIE - Programmable Logic (PL) design to optimize computational efficiency. Our approach explores the parallelization over the rows of the matrices by utilizing as many of the AIE vectorized processors effectively computing all the elements of the resulting vector at the same time, an alternative to cascade stream pipelining. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_15626 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine Sapkas, M. Triossi, A. Zanetti, M. Performance This work explores the use of the AMD Xilinx Versal Adaptable Intelligent Engine (AIE) to accelerate Gated Recurrent Unit (GRU) inference for latency constrained applications. We present a custom workload distribution framework across the AIE's vector processors and propose a hybrid AIE - Programmable Logic (PL) design to optimize computational efficiency. Our approach explores the parallelization over the rows of the matrices by utilizing as many of the AIE vectorized processors effectively computing all the elements of the resulting vector at the same time, an alternative to cascade stream pipelining. |
| title | A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine |
| topic | Performance |
| url | https://arxiv.org/abs/2511.15626 |