DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Chenhao, Ren, Yi, Zhou, Jinwei, Li, Kezhi, Zhang, Chen, Lin, Yibo, Zhang, Lining, Xu, Qiang, Sun, Guangyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-Driven Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025)
von: Ren, Yi, et al.
Veröffentlicht: (2025)
AC-Refiner: Efficient Arithmetic Circuit Optimization Using Conditional Diffusion Models
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
von: Xue, Chenhao, et al.
Veröffentlicht: (2025)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025)
von: Ren, Yi, et al.
Veröffentlicht: (2025)
Design of a Reformed Array Logic Binary Multiplier for High-Speed Computations
von: Mohammad, Sakib, et al.
Veröffentlicht: (2024)
von: Mohammad, Sakib, et al.
Veröffentlicht: (2024)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
Efficient FIR filtering with Bit Layer Multiply Accumulator
von: Liguori, Vincenzo
Veröffentlicht: (2024)
von: Liguori, Vincenzo
Veröffentlicht: (2024)
CellE: Automated Standard Cell Library Extension via Equality Saturation
von: Ren, Yi, et al.
Veröffentlicht: (2026)
von: Ren, Yi, et al.
Veröffentlicht: (2026)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
Efficient Multi-Cycle Folded Integer Multipliers
von: Houraniah, Ahmad, et al.
Veröffentlicht: (2023)
von: Houraniah, Ahmad, et al.
Veröffentlicht: (2023)
Leveraging Highly Approximated Multipliers in DNN Inference
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
An Energy-Efficient Approximate Posit Multiply-Divide Unit
von: Thotli, Rishi, et al.
Veröffentlicht: (2026)
von: Thotli, Rishi, et al.
Veröffentlicht: (2026)
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
von: Liu, Ao, et al.
Veröffentlicht: (2024)
von: Liu, Ao, et al.
Veröffentlicht: (2024)
HPR-Mul: An Area and Energy-Efficient High-Precision Redundancy Multiplier by Approximate Computing
von: Vafaei, Jafar, et al.
Veröffentlicht: (2024)
von: Vafaei, Jafar, et al.
Veröffentlicht: (2024)
On the Systematic Creation of Faithfully Rounded Commutative Truncated Booth Multipliers
von: Drane, Theo, et al.
Veröffentlicht: (2024)
von: Drane, Theo, et al.
Veröffentlicht: (2024)
Analyzing the capabilities of HLS and RTL tools in the design of an FPGA Montgomery Multiplier
von: Ifrim, Rares, et al.
Veröffentlicht: (2025)
von: Ifrim, Rares, et al.
Veröffentlicht: (2025)
Count2Multiply: Reliable In-Memory High-Radix Counting
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2024)
von: de Lima, João Paulo Cardoso, et al.
Veröffentlicht: (2024)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
A Novel FPGA-based CNN Hardware Accelerator: Optimization for Convolutional Layers using Karatsuba Ofman Multiplier
von: Sarkar, Amit
Veröffentlicht: (2024)
von: Sarkar, Amit
Veröffentlicht: (2024)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
von: Gowda, Bindu G, et al.
Veröffentlicht: (2025)
von: Gowda, Bindu G, et al.
Veröffentlicht: (2025)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2024)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
von: Kida, Misaki, et al.
Veröffentlicht: (2025)
von: Kida, Misaki, et al.
Veröffentlicht: (2025)
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
von: Xie, Peichen, et al.
Veröffentlicht: (2025)
von: Xie, Peichen, et al.
Veröffentlicht: (2025)
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
von: Fu, Zizhuo, et al.
Veröffentlicht: (2026)
AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
von: Xue, Chenhao, et al.
Veröffentlicht: (2026)
FPGA-Based Multiplier with a New Approximate Full Adder for Error-Resilient Applications
von: Ranjbar, Ali, et al.
Veröffentlicht: (2025)
von: Ranjbar, Ali, et al.
Veröffentlicht: (2025)
A Logic-Reuse Approach to Nibble-based Multiplier Design for Low Power Vector Computing
von: Chowdhury, Md Rownak Hossain, et al.
Veröffentlicht: (2026)
von: Chowdhury, Md Rownak Hossain, et al.
Veröffentlicht: (2026)
Multiplier Design Addressing Area-Delay Trade-offs by using DSP and Logic resources on FPGAs
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
von: Böttcher, Andreas, et al.
Veröffentlicht: (2024)
Implementation of a 8-bit Wallace Tree Multiplier
von: Biswas, Ayan, et al.
Veröffentlicht: (2025)
von: Biswas, Ayan, et al.
Veröffentlicht: (2025)
AXON: An Automated Netlist Optimization Framework for High-Speed Adders
von: Yang, Tiantian, et al.
Veröffentlicht: (2026)
von: Yang, Tiantian, et al.
Veröffentlicht: (2026)
Approximate Multiplier Induced Error Propagation in Deep Neural Networks
von: Alahakoon, A. M. H. H., et al.
Veröffentlicht: (2025)
von: Alahakoon, A. M. H. H., et al.
Veröffentlicht: (2025)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
von: Jaswal, Pragun, et al.
Veröffentlicht: (2025)
von: Jaswal, Pragun, et al.
Veröffentlicht: (2025)
RL-MUL 2.0: Multiplier Design Optimization with Parallel Deep Reinforcement Learning and Space Reduction
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
Building Reliable Arithmetic Multipliers Under NBTI Aging and Process Variations
von: Heidary, Masoud, et al.
Veröffentlicht: (2026)
von: Heidary, Masoud, et al.
Veröffentlicht: (2026)
Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language Models
von: Zhu, Jingchen, et al.
Veröffentlicht: (2024)
von: Zhu, Jingchen, et al.
Veröffentlicht: (2024)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
A Reconfigurable Multiplier Architecture for Error-Resilient Applications in RISC-V Core
von: Jaswal, Pragun, et al.
Veröffentlicht: (2026)
von: Jaswal, Pragun, et al.
Veröffentlicht: (2026)
Efficient Hardware Implementation of Modular Multiplier over GF (2m) on FPGA
von: Kumari, Ruby, et al.
Veröffentlicht: (2025)
von: Kumari, Ruby, et al.
Veröffentlicht: (2025)
DAISM: Digital Approximate In-SRAM Multiplier-based Accelerator for DNN Training and Inference
von: Sonnino, Lorenzo, et al.
Veröffentlicht: (2023)
von: Sonnino, Lorenzo, et al.
Veröffentlicht: (2023)
A Full-Stack Performance Evaluation Infrastructure for 3D-DRAM-based LLM Accelerators
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DiffuSE: Cross-Layer Design Space Exploration of DNN Accelerator via Diffusion-Driven Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025) -
AC-Refiner: Efficient Arithmetic Circuit Optimization Using Conditional Diffusion Models
von: Xue, Chenhao, et al.
Veröffentlicht: (2025) -
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024) -
Orthrus: Dual-Loop Automated Framework for System-Technology Co-Optimization
von: Ren, Yi, et al.
Veröffentlicht: (2025) -
Design of a Reformed Array Logic Binary Multiplier for High-Speed Computations
von: Mohammad, Sakib, et al.
Veröffentlicht: (2024)