Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
Fuente:
arXiv
Saved in:
| Main Authors: | Gowda, Bindu G, Goyal, Yogesh, Gupta, Yash, Rao, Madhav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HPR-Mul: An Area and Energy-Efficient High-Precision Redundancy Multiplier by Approximate Computing
by: Vafaei, Jafar, et al.
Published: (2024)
by: Vafaei, Jafar, et al.
Published: (2024)
An Energy-Efficient Approximate Posit Multiply-Divide Unit
by: Thotli, Rishi, et al.
Published: (2026)
by: Thotli, Rishi, et al.
Published: (2026)
Hardware Efficient Approximate Convolution with Tunable Error Tolerance for CNNs
by: Shashidhar, Vishal, et al.
Published: (2026)
by: Shashidhar, Vishal, et al.
Published: (2026)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
E2AFS: Energy-Efficient Approximate Floating Point Square Rooter for Error Tolerant Computing
by: Goyal, Prateek, et al.
Published: (2026)
by: Goyal, Prateek, et al.
Published: (2026)
In-Memory Computing Architecture for Efficient Hardware Security
by: Ajmi, Hala, et al.
Published: (2024)
by: Ajmi, Hala, et al.
Published: (2024)
Efficient Hardware Implementation of Modular Multiplier over GF (2m) on FPGA
by: Kumari, Ruby, et al.
Published: (2025)
by: Kumari, Ruby, et al.
Published: (2025)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
by: Wang, Miaoxin, et al.
Published: (2024)
by: Wang, Miaoxin, et al.
Published: (2024)
Efficient Multi-Cycle Folded Integer Multipliers
by: Houraniah, Ahmad, et al.
Published: (2023)
by: Houraniah, Ahmad, et al.
Published: (2023)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
A Novel FPGA-based CNN Hardware Accelerator: Optimization for Convolutional Layers using Karatsuba Ofman Multiplier
by: Sarkar, Amit
Published: (2024)
by: Sarkar, Amit
Published: (2024)
FPGA-Based Multiplier with a New Approximate Full Adder for Error-Resilient Applications
by: Ranjbar, Ali, et al.
Published: (2025)
by: Ranjbar, Ali, et al.
Published: (2025)
ApproXAI: Energy-Efficient Hardware Acceleration of Explainable AI using Approximate Computing
by: Siddique, Ayesha, et al.
Published: (2025)
by: Siddique, Ayesha, et al.
Published: (2025)
Kernel Approximation using Analog In-Memory Computing
by: Büchel, Julian, et al.
Published: (2024)
by: Büchel, Julian, et al.
Published: (2024)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
MixFP4: Enhancing NVFP4 with Adaptive FP4/INT4 Block Representations
by: Zou, Jiaxiang, et al.
Published: (2026)
by: Zou, Jiaxiang, et al.
Published: (2026)
Leveraging Highly Approximated Multipliers in DNN Inference
by: Zervakis, Georgios, et al.
Published: (2024)
by: Zervakis, Georgios, et al.
Published: (2024)
Hardware Efficient Accelerator for Spiking Transformer With Reconfigurable Parallel Time Step Computing
by: Chen, Bo-Yu, et al.
Published: (2025)
by: Chen, Bo-Yu, et al.
Published: (2025)
Design of a Reformed Array Logic Binary Multiplier for High-Speed Computations
by: Mohammad, Sakib, et al.
Published: (2024)
by: Mohammad, Sakib, et al.
Published: (2024)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
by: Zuo, Dongsheng, et al.
Published: (2024)
by: Zuo, Dongsheng, et al.
Published: (2024)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
A Logic-Reuse Approach to Nibble-based Multiplier Design for Low Power Vector Computing
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026)
Design and Analysis of Approximate Hardware Accelerators for VVC Intra Angular Prediction
by: de Fraga, Lucas M. Leipnitz, et al.
Published: (2025)
by: de Fraga, Lucas M. Leipnitz, et al.
Published: (2025)
Approximate Multiplier Induced Error Propagation in Deep Neural Networks
by: Alahakoon, A. M. H. H., et al.
Published: (2025)
by: Alahakoon, A. M. H. H., et al.
Published: (2025)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
by: Zhao, Liang, et al.
Published: (2026)
by: Zhao, Liang, et al.
Published: (2026)
Faster Inference of LLMs using FP8 on the Intel Gaudi
by: Lee, Joonhyung, et al.
Published: (2025)
by: Lee, Joonhyung, et al.
Published: (2025)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
by: Noh, Seock-Hwan, et al.
Published: (2025)
by: Noh, Seock-Hwan, et al.
Published: (2025)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
Generalized Methodology for Determining Numerical Features of Hardware Floating-Point Matrix Multipliers: Part I
by: Khattak, Faizan A, et al.
Published: (2025)
by: Khattak, Faizan A, et al.
Published: (2025)
Efficient and Lightweight In-memory Computing Architecture for Hardware Security
by: Ajmi, Hala, et al.
Published: (2022)
by: Ajmi, Hala, et al.
Published: (2022)
Look-Up Table based Neural Network Hardware
by: Sen, Ovishake, et al.
Published: (2024)
by: Sen, Ovishake, et al.
Published: (2024)
BinSparX: Sparsified Binary Neural Networks for Reduced Hardware Non-Idealities in Xbar Arrays
by: Malhotra, Akul, et al.
Published: (2024)
by: Malhotra, Akul, et al.
Published: (2024)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
by: Li, Zhengke, et al.
Published: (2025)
by: Li, Zhengke, et al.
Published: (2025)
A Power-Efficient Hardware Implementation of L-Mul
by: Chen, Ruiqi, et al.
Published: (2024)
by: Chen, Ruiqi, et al.
Published: (2024)
Register Aggregation for Hardware Decompilation
by: Rao, Varun, et al.
Published: (2024)
by: Rao, Varun, et al.
Published: (2024)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
by: Sonnino, Lorenzo, et al.
Published: (2023)
by: Sonnino, Lorenzo, et al.
Published: (2023)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
by: Wang, Chuanzhen, et al.
Published: (2026)
by: Wang, Chuanzhen, et al.
Published: (2026)
Similar Items
-
HPR-Mul: An Area and Energy-Efficient High-Precision Redundancy Multiplier by Approximate Computing
by: Vafaei, Jafar, et al.
Published: (2024) -
An Energy-Efficient Approximate Posit Multiply-Divide Unit
by: Thotli, Rishi, et al.
Published: (2026) -
Hardware Efficient Approximate Convolution with Tunable Error Tolerance for CNNs
by: Shashidhar, Vishal, et al.
Published: (2026) -
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025) -
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
by: Liu, Ao, et al.
Published: (2024)