Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nair, Harideep, Vellaisamy, Prabhu, Lin, Tsung-Han, Wang, Perry, Blanton, Shawn, Shen, John Paul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
von: Venkatachalam, Shanmuga, et al.
Veröffentlicht: (2026)
von: Venkatachalam, Shanmuga, et al.
Veröffentlicht: (2026)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
von: Nair, Harideep, et al.
Veröffentlicht: (2024)
von: Nair, Harideep, et al.
Veröffentlicht: (2024)
Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks
von: Lister, Devon, et al.
Veröffentlicht: (2025)
von: Lister, Devon, et al.
Veröffentlicht: (2025)
NeRTCAM: CAM-Based CMOS Implementation of Reference Frames for Neuromorphic Processors
von: Nair, Harideep, et al.
Veröffentlicht: (2024)
von: Nair, Harideep, et al.
Veröffentlicht: (2024)
Mugi: Value Level Parallelism For Efficient LLMs
von: Price, Daniel, et al.
Veröffentlicht: (2026)
von: Price, Daniel, et al.
Veröffentlicht: (2026)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
von: Yu, Feng, et al.
Veröffentlicht: (2026)
von: Yu, Feng, et al.
Veröffentlicht: (2026)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
FireFly-T: High-Throughput Sparsity Exploitation for Spiking Transformer Acceleration with Dual-Engine Overlay Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2025)
von: Li, Tenglong, et al.
Veröffentlicht: (2025)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
A Bit Level Weight Reordering Strategy Based on Column Similarity to Explore Weight Sparsity in RRAM-based NN Accelerator
von: Yang, Weiping, et al.
Veröffentlicht: (2025)
von: Yang, Weiping, et al.
Veröffentlicht: (2025)
BBS: Bi-directional Bit-level Sparsity for Deep Learning Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
3D MPSoC with On-Chip Cache Support -- Design and Exploitation
von: Cataldo, Rodrigo, et al.
Veröffentlicht: (2025)
von: Cataldo, Rodrigo, et al.
Veröffentlicht: (2025)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
BitParticle: Partializing Sparse Dual-Factors to Build Quasi-Synchronizing MAC Arrays for Energy-efficient DNNs
von: Qiaoyuan, Feilong, et al.
Veröffentlicht: (2025)
von: Qiaoyuan, Feilong, et al.
Veröffentlicht: (2025)
SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
von: He, Zicheng, et al.
Veröffentlicht: (2026)
von: He, Zicheng, et al.
Veröffentlicht: (2026)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
von: Huang, Zhirui, et al.
Veröffentlicht: (2025)
von: Huang, Zhirui, et al.
Veröffentlicht: (2025)
Weight Transformations in Bit-Sliced Crossbar Arrays for Fault Tolerant Computing-in-Memory: Design Techniques and Evaluation Framework
von: Malhotra, Akul, et al.
Veröffentlicht: (2025)
von: Malhotra, Akul, et al.
Veröffentlicht: (2025)
ADE-HGNN: Accelerating HGNNs through Attention Disparity Exploitation
von: Han, Dengke, et al.
Veröffentlicht: (2024)
von: Han, Dengke, et al.
Veröffentlicht: (2024)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
von: Xie, Rui, et al.
Veröffentlicht: (2025)
von: Xie, Rui, et al.
Veröffentlicht: (2025)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
von: Wei, Jianyu, et al.
Veröffentlicht: (2025)
von: Wei, Jianyu, et al.
Veröffentlicht: (2025)
HiHGNN: Accelerating HGNNs through Parallelism and Data Reusability Exploitation
von: Xue, Runzhen, et al.
Veröffentlicht: (2023)
von: Xue, Runzhen, et al.
Veröffentlicht: (2023)
DL-PIM: Improving Data Locality in Processing-in-Memory Systems
von: Tian, Parker Hao, et al.
Veröffentlicht: (2025)
von: Tian, Parker Hao, et al.
Veröffentlicht: (2025)
Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI Hardware
von: Kanda, R., et al.
Veröffentlicht: (2026)
von: Kanda, R., et al.
Veröffentlicht: (2026)
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
von: Zheng, Keran, et al.
Veröffentlicht: (2025)
von: Zheng, Keran, et al.
Veröffentlicht: (2025)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
von: Choi, Yuseon, et al.
Veröffentlicht: (2025)
Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice Sparsity
von: Kam, Dongyun, et al.
Veröffentlicht: (2024)
von: Kam, Dongyun, et al.
Veröffentlicht: (2024)
Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
von: Gerlinghoff, Daniel, et al.
Veröffentlicht: (2024)
von: Gerlinghoff, Daniel, et al.
Veröffentlicht: (2024)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
Dynamic Power Control in a Hardware Neural Network with Error-Configurable MAC Units
von: Ghaderi, Maedeh, et al.
Veröffentlicht: (2024)
von: Ghaderi, Maedeh, et al.
Veröffentlicht: (2024)
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
von: Zhang, Wenlun, et al.
Veröffentlicht: (2025)
FireFly-S: Exploiting Dual-Side Sparsity for Spiking Neural Networks Acceleration with Reconfigurable Spatial Architecture
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
von: Li, Tenglong, et al.
Veröffentlicht: (2024)
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2026) -
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024) -
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024) -
TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2024) -
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
von: Venkatachalam, Shanmuga, et al.
Veröffentlicht: (2026)