BitParticle: Partializing Sparse Dual-Factors to Build Quasi-Synchronizing MAC Arrays for Energy-efficient DNNs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qiaoyuan, Feilong, Wang, Jihe, Sun, Zhiyu, Wu, Linying, Xiao, Yuanhua, Wang, Danghui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
von: Nair, Harideep, et al.
Veröffentlicht: (2024)
von: Nair, Harideep, et al.
Veröffentlicht: (2024)
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
von: Yu, Feng, et al.
Veröffentlicht: (2026)
von: Yu, Feng, et al.
Veröffentlicht: (2026)
An ECC-based Fault Tolerance Approach for DNNs
von: Raji, Mohsen, et al.
Veröffentlicht: (2025)
von: Raji, Mohsen, et al.
Veröffentlicht: (2025)
Energy-Efficient p-Bit-Based Fully-Connected Quantum-Inspired Simulated Annealer with Dual BRAM Architecture
von: Onizawa, Naoya, et al.
Veröffentlicht: (2026)
von: Onizawa, Naoya, et al.
Veröffentlicht: (2026)
Error Checking for Sparse Systolic Tensor Arrays
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
4T2R X-ReRAM CiM Array for Variation-tolerant, Low-power, Massively Parallel MAC Operation
von: Kihara, Fuyuki, et al.
Veröffentlicht: (2025)
von: Kihara, Fuyuki, et al.
Veröffentlicht: (2025)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
Compromising the Intelligence of Modern DNNs: On the Effectiveness of Targeted RowPress
von: Zhou, Ranyang, et al.
Veröffentlicht: (2024)
von: Zhou, Ranyang, et al.
Veröffentlicht: (2024)
HW-SW Optimization of DNNs for Privacy-preserving People Counting on Low-resolution Infrared Arrays
von: Risso, Matteo, et al.
Veröffentlicht: (2024)
von: Risso, Matteo, et al.
Veröffentlicht: (2024)
ARAS: An Adaptive Low-Cost ReRAM-Based Accelerator for DNNs
von: Sabri, Mohammad, et al.
Veröffentlicht: (2024)
von: Sabri, Mohammad, et al.
Veröffentlicht: (2024)
RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNs
von: Ko, Hanum, et al.
Veröffentlicht: (2026)
von: Ko, Hanum, et al.
Veröffentlicht: (2026)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
Weight Transformations in Bit-Sliced Crossbar Arrays for Fault Tolerant Computing-in-Memory: Design Techniques and Evaluation Framework
von: Malhotra, Akul, et al.
Veröffentlicht: (2025)
von: Malhotra, Akul, et al.
Veröffentlicht: (2025)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
von: Taka, Endri, et al.
Veröffentlicht: (2025)
von: Taka, Endri, et al.
Veröffentlicht: (2025)
Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators
von: Symons, Arne, et al.
Veröffentlicht: (2022)
von: Symons, Arne, et al.
Veröffentlicht: (2022)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
von: Dai, Xilai, et al.
Veröffentlicht: (2024)
von: Dai, Xilai, et al.
Veröffentlicht: (2024)
WAGONN: Weight Bit Agglomeration in Crossbar Arrays for Reduced Impact of Interconnect Resistance on DNN Inference Accuracy
von: Victor, Jeffry, et al.
Veröffentlicht: (2024)
von: Victor, Jeffry, et al.
Veröffentlicht: (2024)
Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
von: Duan, Cenlin, et al.
Veröffentlicht: (2024)
Efficient Reprogramming of Memristive Crossbars for DNNs: Weight Sorting and Bit Stucking
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
Flexible Bit-Truncation Memory for Approximate Applications on the Edge
von: Oswald, William, et al.
Veröffentlicht: (2025)
von: Oswald, William, et al.
Veröffentlicht: (2025)
AccelSync: Verifying Synchronization Coverage in Accelerator Pipeline Programs
von: An, Hangcheng, et al.
Veröffentlicht: (2026)
von: An, Hangcheng, et al.
Veröffentlicht: (2026)
Efficient SRAM-PIM Co-design by Joint Exploration of Value-Level and Bit-Level Sparsity
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
von: Duan, Cenlin, et al.
Veröffentlicht: (2025)
Table-Lookup MAC: Scalable Processing of Quantised Neural Networks in FPGA Soft Logic
von: Gerlinghoff, Daniel, et al.
Veröffentlicht: (2024)
von: Gerlinghoff, Daniel, et al.
Veröffentlicht: (2024)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
von: Zuo, Dongsheng, et al.
Veröffentlicht: (2024)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
von: Ali, Sami Ben, et al.
Veröffentlicht: (2024)
Dynamic Power Control in a Hardware Neural Network with Error-Configurable MAC Units
von: Ghaderi, Maedeh, et al.
Veröffentlicht: (2024)
von: Ghaderi, Maedeh, et al.
Veröffentlicht: (2024)
Sorted Weight Sectioning for Energy-Efficient Unstructured Sparse DNNs on Compute-in-Memory Crossbars
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
von: Farias, Matheus, et al.
Veröffentlicht: (2024)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
von: Yoon, Hyunsung, et al.
Veröffentlicht: (2026)
von: Yoon, Hyunsung, et al.
Veröffentlicht: (2026)
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
von: Wang, Huizheng, et al.
Veröffentlicht: (2025)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
von: Sabih, Muhammad, et al.
Veröffentlicht: (2025)
von: Sabih, Muhammad, et al.
Veröffentlicht: (2025)
Row-Column Hybrid Grouping for Fault-Resilient Multi-Bit Weight Representation on IMC Arrays
von: Jeon, Kang Eun, et al.
Veröffentlicht: (2025)
von: Jeon, Kang Eun, et al.
Veröffentlicht: (2025)
PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
von: Moitra, Abhishek, et al.
Veröffentlicht: (2024)
No One-Size-Fits-All: A Workload-Driven Characterization of Bit-Parallel vs. Bit-Serial Data Layouts for Processing-using-Memory
von: Zhang, Jingyao, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyao, et al.
Veröffentlicht: (2025)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
FractalSync: Lightweight Scalable Global Synchronization of Massive Bulk Synchronous Parallel AI Accelerators
von: Isachi, Victor, et al.
Veröffentlicht: (2025)
von: Isachi, Victor, et al.
Veröffentlicht: (2025)
Periodic Online Testing for Sparse Systolic Tensor Arrays
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2025)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2025)
A 64-Spin All-to-All CMOS Ising Machine with Landscape Perturbation Achieving 2.28 nJ/Edge-Bit Energy-to-Solution
von: Salim, Ahmet Yusuf, et al.
Veröffentlicht: (2026)
von: Salim, Ahmet Yusuf, et al.
Veröffentlicht: (2026)
Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs
von: Ahmadilivani, Mohammad Hasan, et al.
Veröffentlicht: (2026)
von: Ahmadilivani, Mohammad Hasan, et al.
Veröffentlicht: (2026)
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
von: Li, Zhengke, et al.
Veröffentlicht: (2025)
von: Li, Zhengke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
von: Nair, Harideep, et al.
Veröffentlicht: (2024) -
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA
von: Yu, Feng, et al.
Veröffentlicht: (2026) -
An ECC-based Fault Tolerance Approach for DNNs
von: Raji, Mohsen, et al.
Veröffentlicht: (2025) -
Energy-Efficient p-Bit-Based Fully-Connected Quantum-Inspired Simulated Annealer with Dual BRAM Architecture
von: Onizawa, Naoya, et al.
Veröffentlicht: (2026) -
Error Checking for Sparse Systolic Tensor Arrays
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)