CORDIC Is All You Need

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kokane, Omkar, Teman, Adam, Jha, Anushka, SL, Guru Prasath, Raut, Gopal, Lokhande, Mukul, Chand, S. V. Jaya, Dewangan, Tanushree, Vishvakarma, Santosh Kumar
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917958065848320
author Kokane, Omkar
Teman, Adam
Jha, Anushka
SL, Guru Prasath
Raut, Gopal
Lokhande, Mukul
Chand, S. V. Jaya
Dewangan, Tanushree
Vishvakarma, Santosh Kumar
author_facet Kokane, Omkar
Teman, Adam
Jha, Anushka
SL, Guru Prasath
Raut, Gopal
Lokhande, Mukul
Chand, S. V. Jaya
Dewangan, Tanushree
Vishvakarma, Santosh Kumar
contents Artificial intelligence necessitates adaptable hardware accelerators for efficient high-throughput million operations. We present pipelined architecture with CORDIC block for linear MAC computations and nonlinear iterative Activation Functions (AF) such as $tanh$, $sigmoid$, and $softmax$. This approach focuses on a Reconfigurable Processing Engine (RPE) based systolic array, with 40\% pruning rate, enhanced throughput up to 4.64$\times$, and reduction in power and area by 5.02 $\times$ and 4.06 $\times$ at CMOS 28 nm, with minor accuracy loss. FPGA implementation achieves a reduction of up to 2.5 $\times$ resource savings and 3 $\times$ power compared to prior works. The Systolic CORDIC engine for Reconfigurability and Enhanced throughput (SYCore) deploys an output stationary dataflow with the CAESAR control engine for diverse AI workloads such as Transformers, RNNs/LSTMs, and DNNs for applications like image detection, LLMs, and speech recognition. The energy-efficient and flexible approach extends the enhanced approach for edge AI accelerators supporting emerging workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11685
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CORDIC Is All You Need
Kokane, Omkar
Teman, Adam
Jha, Anushka
SL, Guru Prasath
Raut, Gopal
Lokhande, Mukul
Chand, S. V. Jaya
Dewangan, Tanushree
Vishvakarma, Santosh Kumar
Hardware Architecture
Computer Vision and Pattern Recognition
Image and Video Processing
Artificial intelligence necessitates adaptable hardware accelerators for efficient high-throughput million operations. We present pipelined architecture with CORDIC block for linear MAC computations and nonlinear iterative Activation Functions (AF) such as $tanh$, $sigmoid$, and $softmax$. This approach focuses on a Reconfigurable Processing Engine (RPE) based systolic array, with 40\% pruning rate, enhanced throughput up to 4.64$\times$, and reduction in power and area by 5.02 $\times$ and 4.06 $\times$ at CMOS 28 nm, with minor accuracy loss. FPGA implementation achieves a reduction of up to 2.5 $\times$ resource savings and 3 $\times$ power compared to prior works. The Systolic CORDIC engine for Reconfigurability and Enhanced throughput (SYCore) deploys an output stationary dataflow with the CAESAR control engine for diverse AI workloads such as Transformers, RNNs/LSTMs, and DNNs for applications like image detection, LLMs, and speech recognition. The energy-efficient and flexible approach extends the enhanced approach for edge AI accelerators supporting emerging workloads.
title CORDIC Is All You Need
topic Hardware Architecture
Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2503.11685