SRAM Based Digital Custom Compute Engine for Improved Area Efficiency of AI Hardware

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dhakad, Narendra Singh, Vishvakarma, Santosh Kumar
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911688701247488
author Dhakad, Narendra Singh
Vishvakarma, Santosh Kumar
author_facet Dhakad, Narendra Singh
Vishvakarma, Santosh Kumar
contents This paper presents a novel architecture utilizing a 10T SRAM cell for XNOR-based in-memory computing, aimed at mitigating the extensive routing challenges typically encountered in conventional in-memory computing systems. By integrating a full adder between in-memory multiplication cells, the proposed design achieves a 50% reduction in routing complexity. The architecture performs multiply-accumulate (MAC) operations using XNOR computation optimized for binary neural networks (BNNs). Additionally, a 14T-based full adder is employed to construct an N-bit ripple carry adder in the adder tree, significantly reducing the area compared to traditional 28T-based CMOS designs. The 10T SRAM XNOR computation further enhances the latency for MAC operations. The proposed approach reduces the latency and area overhead, improving the overall hardware's area efficiency by 2.67x compared to the state-of-the-art.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16161
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SRAM Based Digital Custom Compute Engine for Improved Area Efficiency of AI Hardware
Dhakad, Narendra Singh
Vishvakarma, Santosh Kumar
Hardware Architecture
This paper presents a novel architecture utilizing a 10T SRAM cell for XNOR-based in-memory computing, aimed at mitigating the extensive routing challenges typically encountered in conventional in-memory computing systems. By integrating a full adder between in-memory multiplication cells, the proposed design achieves a 50% reduction in routing complexity. The architecture performs multiply-accumulate (MAC) operations using XNOR computation optimized for binary neural networks (BNNs). Additionally, a 14T-based full adder is employed to construct an N-bit ripple carry adder in the adder tree, significantly reducing the area compared to traditional 28T-based CMOS designs. The 10T SRAM XNOR computation further enhances the latency for MAC operations. The proposed approach reduces the latency and area overhead, improving the overall hardware's area efficiency by 2.67x compared to the state-of-the-art.
title SRAM Based Digital Custom Compute Engine for Improved Area Efficiency of AI Hardware
topic Hardware Architecture
url https://arxiv.org/abs/2605.16161