bitSMM: A bit-Serial Matrix Multiplication Accelerator
Fuente:
arXiv
Saved in:
| Main Authors: | Antunes, Pedro, Podobas, Artur |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FPGA-Based Neural Network Accelerators for Space Applications: A Survey
by: Antunes, Pedro, et al.
Published: (2025)
by: Antunes, Pedro, et al.
Published: (2025)
A Quarter of a Century of Neuromorphic Architectures on FPGAs -- an Overview
by: Szczerek, Wiktor J., et al.
Published: (2025)
by: Szczerek, Wiktor J., et al.
Published: (2025)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
IzhiRISC-V -- a RISC-V-based Processor with Custom ISA Extension for Spiking Neuron Networks Processing with Izhikevich Neurons
by: Szczerek, Wiktor J., et al.
Published: (2025)
by: Szczerek, Wiktor J., et al.
Published: (2025)
Fast Algorithms for Spiking Neural Network Simulation with FPGAs
by: Lindqvist, Björn A., et al.
Published: (2024)
by: Lindqvist, Björn A., et al.
Published: (2024)
Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2025)
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2025)
Evaluating Four FPGA-accelerated Space Use Cases based on Neural Network Algorithms for On-board Inference
by: Antunes, Pedro, et al.
Published: (2026)
by: Antunes, Pedro, et al.
Published: (2026)
Hardware-Efficient Accurate 4-bit Multiplier for Xilinx 7 Series FPGAs
by: Kida, Misaki, et al.
Published: (2025)
by: Kida, Misaki, et al.
Published: (2025)
A Reconfigurable Stream-Based FPGA Accelerator for Bayesian Confidence Propagation Neural Networks
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2025)
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2025)
A 33.6-136.2 TOPS/W Nonlinear Analog Computing-In-Memory Macro for Multi-bit LSTM Accelerator in 65 nm CMOS
by: Yang, Junyi, et al.
Published: (2025)
by: Yang, Junyi, et al.
Published: (2025)
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
by: Hu, Weiming, et al.
Published: (2026)
by: Hu, Weiming, et al.
Published: (2026)
WebRISC-V: A 64-bit RISC-V Pipeline Simulator for Computer Architecture Classes
by: Giorgi, Roberto, et al.
Published: (2025)
by: Giorgi, Roberto, et al.
Published: (2025)
Implementation of a 8-bit Wallace Tree Multiplier
by: Biswas, Ayan, et al.
Published: (2025)
by: Biswas, Ayan, et al.
Published: (2025)
Optimization of 32-bit Unsigned Division by Constants on 64-bit Targets
by: Mitsunari, Shigeo, et al.
Published: (2026)
by: Mitsunari, Shigeo, et al.
Published: (2026)
Modulo-$(2^{2n}+1)$ Arithmetic via Two Parallel n-bit Residue Channels
by: Jaberipur, Ghassem, et al.
Published: (2024)
by: Jaberipur, Ghassem, et al.
Published: (2024)
Semicustom Frontend VLSI Design and Analysis of a 32-bit Brent-Kung Adder in Cadence Suite
by: Singh, Yashvardhan
Published: (2025)
by: Singh, Yashvardhan
Published: (2025)
M-ANT: Efficient Low-bit Group Quantization for LLMs via Mathematically Adaptive Numerical Type
by: Hu, Weiming, et al.
Published: (2025)
by: Hu, Weiming, et al.
Published: (2025)
Design of a 6-bit Threshold Inverter Quantization (TIQ) Flash Analog to Digital Converter (ADC)
by: Sarkar, Noyon Kumar, et al.
Published: (2025)
by: Sarkar, Noyon Kumar, et al.
Published: (2025)
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
by: Abdelmaksoud, Ahmed J., et al.
Published: (2025)
by: Abdelmaksoud, Ahmed J., et al.
Published: (2025)
On Approximate 8-bit Floating-Point Operations Using Integer Operations
by: Lindberg, Theodor, et al.
Published: (2024)
by: Lindberg, Theodor, et al.
Published: (2024)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
by: Ke, Chih-Hua
Published: (2026)
by: Ke, Chih-Hua
Published: (2026)
'1'-bit Count-based Sorting Unit to Reduce Link Power in DNN Accelerators
by: Han, Ruichi, et al.
Published: (2026)
by: Han, Ruichi, et al.
Published: (2026)
BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
by: Zhang, Wenlun, et al.
Published: (2025)
by: Zhang, Wenlun, et al.
Published: (2025)
GUST: Graph Edge-Coloring Utilization for Accelerating Sparse Matrix Vector Multiplication
by: Gerami, Armin, et al.
Published: (2024)
by: Gerami, Armin, et al.
Published: (2024)
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
by: Abdelmaksoud, Ahmed J., et al.
Published: (2026)
by: Abdelmaksoud, Ahmed J., et al.
Published: (2026)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025)
by: Malekar, Jinendra, et al.
Published: (2025)
KANtize: Exploring Low-bit Quantization of Kolmogorov-Arnold Networks for Efficient Inference
by: Errabii, Sohaib, et al.
Published: (2026)
by: Errabii, Sohaib, et al.
Published: (2026)
GAVINA: flexible aggressive undervolting for bit-serial mixed-precision DNN acceleration
by: Fornt, Jordi, et al.
Published: (2025)
by: Fornt, Jordi, et al.
Published: (2025)
CVA6-VMRT: A Modular Approach Towards Time-Predictable Virtual Memory in a 64-bit Application Class RISC-V Processor
by: Reinwardt, Christopher, et al.
Published: (2025)
by: Reinwardt, Christopher, et al.
Published: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
by: Shan, Haoxuan, et al.
Published: (2025)
by: Shan, Haoxuan, et al.
Published: (2025)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
by: Su, Yuchao, et al.
Published: (2025)
by: Su, Yuchao, et al.
Published: (2025)
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2026)
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2026)
Basilisk: A 34 mm2 End-to-End Open-Source 64-bit Linux-Capable RISC-V SoC in 130nm BiCMOS
by: Sauter, Philippe, et al.
Published: (2025)
by: Sauter, Philippe, et al.
Published: (2025)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
by: Yoon, Hyunsung, et al.
Published: (2026)
by: Yoon, Hyunsung, et al.
Published: (2026)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
High-Performance Pipelined NTT Accelerators with Homogeneous Digit-Serial Modulo Arithmetic
by: Alexakis, George, et al.
Published: (2025)
by: Alexakis, George, et al.
Published: (2025)
DSLR-CNN: Efficient CNN Acceleration using Digit-Serial Left-to-Right Arithmetic
by: Nisar, Malik Zohaib, et al.
Published: (2025)
by: Nisar, Malik Zohaib, et al.
Published: (2025)
FAME: FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption
by: Xu, Zhihan, et al.
Published: (2025)
by: Xu, Zhihan, et al.
Published: (2025)
Q3DE: A fault-tolerant quantum computer architecture for multi-bit burst errors by cosmic rays
by: Suzuki, Yasunari, et al.
Published: (2024)
by: Suzuki, Yasunari, et al.
Published: (2024)
Occamy: A 432-Core Dual-Chiplet Dual-HBM2E 768-DP-GFLOP/s RISC-V System for 8-to-64-bit Dense and Sparse Computing in 12nm FinFET
by: Scheffler, Paul, et al.
Published: (2025)
by: Scheffler, Paul, et al.
Published: (2025)
Similar Items
-
FPGA-Based Neural Network Accelerators for Space Applications: A Survey
by: Antunes, Pedro, et al.
Published: (2025) -
A Quarter of a Century of Neuromorphic Architectures on FPGAs -- an Overview
by: Szczerek, Wiktor J., et al.
Published: (2025) -
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026) -
IzhiRISC-V -- a RISC-V-based Processor with Custom ISA Extension for Spiking Neuron Networks Processing with Izhikevich Neurons
by: Szczerek, Wiktor J., et al.
Published: (2025) -
Fast Algorithms for Spiking Neural Network Simulation with FPGAs
by: Lindqvist, Björn A., et al.
Published: (2024)