tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
Fuente:
arXiv
Saved in:
| Main Authors: | Vellaisamy, Prabhu, Nair, Harideep, Finn, Joseph, Trivedi, Manav, Chen, Albert, Li, Anna, Lin, Tsung-Han, Wang, Perry, Blanton, Shawn, Shen, John Paul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
by: Vellaisamy, Prabhu, et al.
Published: (2026)
by: Vellaisamy, Prabhu, et al.
Published: (2026)
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks
by: Lister, Devon, et al.
Published: (2025)
by: Lister, Devon, et al.
Published: (2025)
Scaling Analog Photonic Accelerators for Byte-Size, Integer General Matrix Multiply (GEMM) Kernels
by: Alo, Oluwaseun Adewunmi, et al.
Published: (2024)
by: Alo, Oluwaseun Adewunmi, et al.
Published: (2024)
Power- and Area-Efficient Unary Sorting Architecture Using FSM-Based Unary Number Generator
by: Jalilvand, Amir Hossein, et al.
Published: (2025)
by: Jalilvand, Amir Hossein, et al.
Published: (2025)
Mugi: Value Level Parallelism For Efficient LLMs
by: Price, Daniel, et al.
Published: (2026)
by: Price, Daniel, et al.
Published: (2026)
An Energy-Efficient Approximate Posit Multiply-Divide Unit
by: Thotli, Rishi, et al.
Published: (2026)
by: Thotli, Rishi, et al.
Published: (2026)
NeRTCAM: CAM-Based CMOS Implementation of Reference Frames for Neuromorphic Processors
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
Scaling Photonic Tensor Cores with Unary and Homodyne Designs
by: Alo, Oluwaseun, et al.
Published: (2026)
by: Alo, Oluwaseun, et al.
Published: (2026)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025)
by: Ta, Tuan, et al.
Published: (2025)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
by: Noh, Seock-Hwan, et al.
Published: (2025)
by: Noh, Seock-Hwan, et al.
Published: (2025)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)
by: Li, Haomin, et al.
Published: (2026)
Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
by: Xie, Peichen, et al.
Published: (2025)
by: Xie, Peichen, et al.
Published: (2025)
Design of a Reformed Array Logic Binary Multiplier for High-Speed Computations
by: Mohammad, Sakib, et al.
Published: (2024)
by: Mohammad, Sakib, et al.
Published: (2024)
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)
by: Papalamprou, Ilias, et al.
Published: (2025)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
by: Guo, Cong, et al.
Published: (2025)
by: Guo, Cong, et al.
Published: (2025)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
PULSE: Parametric Hardware Units for Low-power Sparsity-Aware Convolution Engine
by: Aliyev, Ilkin, et al.
Published: (2024)
by: Aliyev, Ilkin, et al.
Published: (2024)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
MACO: Exploring GEMM Acceleration on a Loosely-Coupled Multi-core Processor
by: Sui, Bingcai, et al.
Published: (2024)
by: Sui, Bingcai, et al.
Published: (2024)
A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
by: Mhatre, Kaustubh, et al.
Published: (2025)
by: Mhatre, Kaustubh, et al.
Published: (2025)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
by: Zuo, Dongsheng, et al.
Published: (2024)
by: Zuo, Dongsheng, et al.
Published: (2024)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Reducing the Cost of Dropout in Flash-Attention by Hiding RNG with GEMM
by: Ma, Haiyue, et al.
Published: (2024)
by: Ma, Haiyue, et al.
Published: (2024)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
by: Cammarata, Danilo, et al.
Published: (2026)
by: Cammarata, Danilo, et al.
Published: (2026)
Generalized Methodology for Determining Numerical Features of Hardware Floating-Point Matrix Multipliers: Part I
by: Khattak, Faizan A, et al.
Published: (2025)
by: Khattak, Faizan A, et al.
Published: (2025)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
by: Su, Zhongling, et al.
Published: (2025)
by: Su, Zhongling, et al.
Published: (2025)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
by: Kim, Hansung, et al.
Published: (2024)
by: Kim, Hansung, et al.
Published: (2024)
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
by: Xue, Chenhao, et al.
Published: (2025)
by: Xue, Chenhao, et al.
Published: (2025)
Efficient Multi-Cycle Folded Integer Multipliers
by: Houraniah, Ahmad, et al.
Published: (2023)
by: Houraniah, Ahmad, et al.
Published: (2023)
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Similar Items
-
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
by: Nair, Harideep, et al.
Published: (2024) -
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
by: Vellaisamy, Prabhu, et al.
Published: (2026) -
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
by: Nair, Harideep, et al.
Published: (2024) -
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
by: Vellaisamy, Prabhu, et al.
Published: (2024) -
TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering
by: Vellaisamy, Prabhu, et al.
Published: (2024)