tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
Fuente:
arXiv
Saved in:
| Main Authors: | Nair, Harideep, Vellaisamy, Prabhu, Chen, Albert, Finn, Joseph, Li, Anna, Trivedi, Manav, Shen, John Paul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
by: Vellaisamy, Prabhu, et al.
Published: (2026)
by: Vellaisamy, Prabhu, et al.
Published: (2026)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks
by: Lister, Devon, et al.
Published: (2025)
by: Lister, Devon, et al.
Published: (2025)
Optimizing GEMM for Energy and Performance on Versal ACAP Architectures
by: Papalamprou, Ilias, et al.
Published: (2025)
by: Papalamprou, Ilias, et al.
Published: (2025)
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
Power- and Area-Efficient Unary Sorting Architecture Using FSM-Based Unary Number Generator
by: Jalilvand, Amir Hossein, et al.
Published: (2025)
by: Jalilvand, Amir Hossein, et al.
Published: (2025)
GEMM-GS: Accelerating 3D Gaussian Splatting on Tensor Cores with GEMM-Compatible Blending
by: Li, Haomin, et al.
Published: (2026)
by: Li, Haomin, et al.
Published: (2026)
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
by: Taka, Endri, et al.
Published: (2024)
by: Taka, Endri, et al.
Published: (2024)
A Scalable FPGA Architecture With Adaptive Memory Utilization for GEMM-Based Operations
by: Petropoulos, Anastasios, et al.
Published: (2025)
by: Petropoulos, Anastasios, et al.
Published: (2025)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
by: Su, Zhongling, et al.
Published: (2025)
by: Su, Zhongling, et al.
Published: (2025)
TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
Transitive Array: An Efficient GEMM Accelerator with Result Reuse
by: Guo, Cong, et al.
Published: (2025)
by: Guo, Cong, et al.
Published: (2025)
MACO: Exploring GEMM Acceleration on a Loosely-Coupled Multi-core Processor
by: Sui, Bingcai, et al.
Published: (2024)
by: Sui, Bingcai, et al.
Published: (2024)
GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
by: Mhatre, Kaustubh, et al.
Published: (2025)
by: Mhatre, Kaustubh, et al.
Published: (2025)
Reducing the Cost of Dropout in Flash-Attention by Hiding RNG with GEMM
by: Ma, Haiyue, et al.
Published: (2024)
by: Ma, Haiyue, et al.
Published: (2024)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
A Low-Dissipation and Scalable GEMM Accelerator with Silicon Nitride Photonics
by: Karempudi, Venkata Sai Praneeth, et al.
Published: (2024)
by: Karempudi, Venkata Sai Praneeth, et al.
Published: (2024)
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
by: Grailoo, M., et al.
Published: (2026)
by: Grailoo, M., et al.
Published: (2026)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
by: Cammarata, Danilo, et al.
Published: (2026)
by: Cammarata, Danilo, et al.
Published: (2026)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
by: Ta, Tuan, et al.
Published: (2025)
by: Ta, Tuan, et al.
Published: (2025)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
by: Shen, Aofeng, et al.
Published: (2025)
by: Shen, Aofeng, et al.
Published: (2025)
Mugi: Value Level Parallelism For Efficient LLMs
by: Price, Daniel, et al.
Published: (2026)
by: Price, Daniel, et al.
Published: (2026)
Scaling Analog Photonic Accelerators for Byte-Size, Integer General Matrix Multiply (GEMM) Kernels
by: Alo, Oluwaseun Adewunmi, et al.
Published: (2024)
by: Alo, Oluwaseun Adewunmi, et al.
Published: (2024)
NeRTCAM: CAM-Based CMOS Implementation of Reference Frames for Neuromorphic Processors
by: Nair, Harideep, et al.
Published: (2024)
by: Nair, Harideep, et al.
Published: (2024)
SISA: A Scale-In Systolic Array for GEMM Acceleration
by: Altamura, Luigi, et al.
Published: (2026)
by: Altamura, Luigi, et al.
Published: (2026)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025)
by: Müller, Mika Markus, et al.
Published: (2025)
A Comparative Analysis of Microrings Based Incoherent Photonic GEMM Accelerators
by: Vatsavai, Sairam Sri, et al.
Published: (2024)
by: Vatsavai, Sairam Sri, et al.
Published: (2024)
Scaling Photonic Tensor Cores with Unary and Homodyne Designs
by: Alo, Oluwaseun, et al.
Published: (2026)
by: Alo, Oluwaseun, et al.
Published: (2026)
MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology
by: Jeong, Minki, et al.
Published: (2022)
by: Jeong, Minki, et al.
Published: (2022)
A Compact, Low Power Transprecision ALU for Smart Edge Devices
by: Dube, Ayushi, et al.
Published: (2025)
by: Dube, Ayushi, et al.
Published: (2025)
PowerFlow-DNN: Compiler-Directed Fine-Grained Power Orchestration for End-to-End Edge AI Inference
by: Chen, Paul, et al.
Published: (2026)
by: Chen, Paul, et al.
Published: (2026)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Enhanced Hybrid Temporal Computing Using Deterministic Summations for Ultra-Low-Power Accelerators
by: Sachdeva, Sachin, et al.
Published: (2025)
by: Sachdeva, Sachin, et al.
Published: (2025)
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
by: Min, Kyeongpil, et al.
Published: (2026)
by: Min, Kyeongpil, et al.
Published: (2026)
Switchable Single/Dual Edge Registers for Pipeline Architecture
by: Singh, Suyash Vardhan, et al.
Published: (2024)
by: Singh, Suyash Vardhan, et al.
Published: (2024)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
Similar Items
-
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
by: Vellaisamy, Prabhu, et al.
Published: (2024) -
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
by: Vellaisamy, Prabhu, et al.
Published: (2024) -
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
by: Vellaisamy, Prabhu, et al.
Published: (2026) -
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
by: Venkatachalam, Shanmuga, et al.
Published: (2026) -
Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks
by: Lister, Devon, et al.
Published: (2025)