CORDIC Is All You Need
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917958065848320 |
|---|---|
| author | Kokane, Omkar Teman, Adam Jha, Anushka SL, Guru Prasath Raut, Gopal Lokhande, Mukul Chand, S. V. Jaya Dewangan, Tanushree Vishvakarma, Santosh Kumar |
| author_facet | Kokane, Omkar Teman, Adam Jha, Anushka SL, Guru Prasath Raut, Gopal Lokhande, Mukul Chand, S. V. Jaya Dewangan, Tanushree Vishvakarma, Santosh Kumar |
| contents | Artificial intelligence necessitates adaptable hardware accelerators for efficient high-throughput million operations. We present pipelined architecture with CORDIC block for linear MAC computations and nonlinear iterative Activation Functions (AF) such as $tanh$, $sigmoid$, and $softmax$. This approach focuses on a Reconfigurable Processing Engine (RPE) based systolic array, with 40\% pruning rate, enhanced throughput up to 4.64$\times$, and reduction in power and area by 5.02 $\times$ and 4.06 $\times$ at CMOS 28 nm, with minor accuracy loss. FPGA implementation achieves a reduction of up to 2.5 $\times$ resource savings and 3 $\times$ power compared to prior works. The Systolic CORDIC engine for Reconfigurability and Enhanced throughput (SYCore) deploys an output stationary dataflow with the CAESAR control engine for diverse AI workloads such as Transformers, RNNs/LSTMs, and DNNs for applications like image detection, LLMs, and speech recognition. The energy-efficient and flexible approach extends the enhanced approach for edge AI accelerators supporting emerging workloads. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_11685 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CORDIC Is All You Need Kokane, Omkar Teman, Adam Jha, Anushka SL, Guru Prasath Raut, Gopal Lokhande, Mukul Chand, S. V. Jaya Dewangan, Tanushree Vishvakarma, Santosh Kumar Hardware Architecture Computer Vision and Pattern Recognition Image and Video Processing Artificial intelligence necessitates adaptable hardware accelerators for efficient high-throughput million operations. We present pipelined architecture with CORDIC block for linear MAC computations and nonlinear iterative Activation Functions (AF) such as $tanh$, $sigmoid$, and $softmax$. This approach focuses on a Reconfigurable Processing Engine (RPE) based systolic array, with 40\% pruning rate, enhanced throughput up to 4.64$\times$, and reduction in power and area by 5.02 $\times$ and 4.06 $\times$ at CMOS 28 nm, with minor accuracy loss. FPGA implementation achieves a reduction of up to 2.5 $\times$ resource savings and 3 $\times$ power compared to prior works. The Systolic CORDIC engine for Reconfigurability and Enhanced throughput (SYCore) deploys an output stationary dataflow with the CAESAR control engine for diverse AI workloads such as Transformers, RNNs/LSTMs, and DNNs for applications like image detection, LLMs, and speech recognition. The energy-efficient and flexible approach extends the enhanced approach for edge AI accelerators supporting emerging workloads. |
| title | CORDIC Is All You Need |
| topic | Hardware Architecture Computer Vision and Pattern Recognition Image and Video Processing |
| url | https://arxiv.org/abs/2503.11685 |