Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
Fuente:
arXiv
Guardado en:
| Autores principales: | Vellaisamy, Prabhu, Nair, Harideep, Wu, Di, Blanton, Shawn, Shen, John Paul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
por: Nair, Harideep, et al.
Publicado: (2024)
por: Nair, Harideep, et al.
Publicado: (2024)
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
por: Nair, Harideep, et al.
Publicado: (2024)
por: Nair, Harideep, et al.
Publicado: (2024)
TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)
Catwalk: Unary Top-K for Efficient Ramp-No-Leak Neuron Design for Temporal Neural Networks
por: Lister, Devon, et al.
Publicado: (2025)
por: Lister, Devon, et al.
Publicado: (2025)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
por: Venkatachalam, Shanmuga, et al.
Publicado: (2026)
por: Venkatachalam, Shanmuga, et al.
Publicado: (2026)
Mugi: Value Level Parallelism For Efficient LLMs
por: Price, Daniel, et al.
Publicado: (2026)
por: Price, Daniel, et al.
Publicado: (2026)
NeRTCAM: CAM-Based CMOS Implementation of Reference Frames for Neuromorphic Processors
por: Nair, Harideep, et al.
Publicado: (2024)
por: Nair, Harideep, et al.
Publicado: (2024)
Building Reliable Arithmetic Multipliers Under NBTI Aging and Process Variations
por: Heidary, Masoud, et al.
Publicado: (2026)
por: Heidary, Masoud, et al.
Publicado: (2026)
Power- and Area-Efficient Unary Sorting Architecture Using FSM-Based Unary Number Generator
por: Jalilvand, Amir Hossein, et al.
Publicado: (2025)
por: Jalilvand, Amir Hossein, et al.
Publicado: (2025)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
por: Wu, Jiajun, et al.
Publicado: (2024)
por: Wu, Jiajun, et al.
Publicado: (2024)
Idle is the New Sleep: Configuration-Aware Alternative to Powering Off FPGA-Based DL Accelerators During Inactivity
por: Qian, Chao, et al.
Publicado: (2024)
por: Qian, Chao, et al.
Publicado: (2024)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
por: Meng, Chang, et al.
Publicado: (2026)
por: Meng, Chang, et al.
Publicado: (2026)
An Energy-Efficient Approximate Posit Multiply-Divide Unit
por: Thotli, Rishi, et al.
Publicado: (2026)
por: Thotli, Rishi, et al.
Publicado: (2026)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
por: Jaswal, Pragun, et al.
Publicado: (2025)
por: Jaswal, Pragun, et al.
Publicado: (2025)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
por: Alexandridis, Kosmas, et al.
Publicado: (2024)
por: Alexandridis, Kosmas, et al.
Publicado: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
por: Vellaisamy, Prabhu, et al.
Publicado: (2025)
YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AI
por: Xuan, Zihao, et al.
Publicado: (2023)
por: Xuan, Zihao, et al.
Publicado: (2023)
Increasing the Energy-Efficiency of Wearables Using Low-Precision Posit Arithmetic with PHEE
por: Mallasén, David, et al.
Publicado: (2025)
por: Mallasén, David, et al.
Publicado: (2025)
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
por: Shukla, Arnav, et al.
Publicado: (2025)
por: Shukla, Arnav, et al.
Publicado: (2025)
Enable Lightweight and Precision-Scalable Posit/IEEE-754 Arithmetic in RISC-V Cores for Transprecision Computing
por: Li, Qiong, et al.
Publicado: (2025)
por: Li, Qiong, et al.
Publicado: (2025)
Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
por: Noh, Seock-Hwan, et al.
Publicado: (2025)
por: Noh, Seock-Hwan, et al.
Publicado: (2025)
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2025)
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2025)
Scaling Photonic Tensor Cores with Unary and Homodyne Designs
por: Alo, Oluwaseun, et al.
Publicado: (2026)
por: Alo, Oluwaseun, et al.
Publicado: (2026)
GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators
por: Liu, Yuhao, et al.
Publicado: (2026)
por: Liu, Yuhao, et al.
Publicado: (2026)
Voyager: An End-to-End Framework for Design-Space Exploration and Generation of DNN Accelerators
por: Prabhu, Kartik, et al.
Publicado: (2025)
por: Prabhu, Kartik, et al.
Publicado: (2025)
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
por: Zhang, Jinsong, et al.
Publicado: (2025)
por: Zhang, Jinsong, et al.
Publicado: (2025)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
por: Shan, Haoxuan, et al.
Publicado: (2025)
por: Shan, Haoxuan, et al.
Publicado: (2025)
AdAM: Adaptive Fault-Tolerant Approximate Multiplier for Edge DNN Accelerators
por: Taheri, Mahdi, et al.
Publicado: (2024)
por: Taheri, Mahdi, et al.
Publicado: (2024)
GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units
por: Bouvier, Maxence, et al.
Publicado: (2025)
por: Bouvier, Maxence, et al.
Publicado: (2025)
HPR-Mul: An Area and Energy-Efficient High-Precision Redundancy Multiplier by Approximate Computing
por: Vafaei, Jafar, et al.
Publicado: (2024)
por: Vafaei, Jafar, et al.
Publicado: (2024)
A Reconfigurable Multiplier Architecture for Error-Resilient Applications in RISC-V Core
por: Jaswal, Pragun, et al.
Publicado: (2026)
por: Jaswal, Pragun, et al.
Publicado: (2026)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
por: Böttcher, Andreas, et al.
Publicado: (2024)
por: Böttcher, Andreas, et al.
Publicado: (2024)
Instruction-Based Coordination of Heterogeneous Processing Units for Acceleration of DNN Inference
por: Petropoulos, Anastasios, et al.
Publicado: (2025)
por: Petropoulos, Anastasios, et al.
Publicado: (2025)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
por: Sharma, Vinamra, et al.
Publicado: (2026)
por: Sharma, Vinamra, et al.
Publicado: (2026)
A Blueprint for Precise and Fault-Tolerant Analog Neural Networks
por: Demirkiran, Cansu, et al.
Publicado: (2023)
por: Demirkiran, Cansu, et al.
Publicado: (2023)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
por: Parthasarathy, Rishab
Publicado: (2025)
por: Parthasarathy, Rishab
Publicado: (2025)
Scaling Analog Photonic Accelerators for Byte-Size, Integer General Matrix Multiply (GEMM) Kernels
por: Alo, Oluwaseun Adewunmi, et al.
Publicado: (2024)
por: Alo, Oluwaseun Adewunmi, et al.
Publicado: (2024)
Multiplier-free In-Memory Vector-Matrix Multiplication Using Distributed Arithmetic
por: Zeller, Felix, et al.
Publicado: (2025)
por: Zeller, Felix, et al.
Publicado: (2025)
Ejemplares similares
-
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
por: Vellaisamy, Prabhu, et al.
Publicado: (2024) -
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
por: Vellaisamy, Prabhu, et al.
Publicado: (2024) -
Commercial Evaluation of Zero-Skipping MAC Design for Bit Sparsity Exploitation in DL Inference
por: Nair, Harideep, et al.
Publicado: (2024) -
tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI
por: Nair, Harideep, et al.
Publicado: (2024) -
TNNGen: Automated Design of Neuromorphic Sensory Processing Units for Time-Series Clustering
por: Vellaisamy, Prabhu, et al.
Publicado: (2024)