Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
Fuente:
arXiv
Saved in:
| Main Authors: | Alexandridis, Kosmas, Peltekis, Christodoulos, Filippas, Dionysios, Dimitrakopoulos, Giorgos |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Error Checking for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Online Alignment and Addition in Multi-Term Floating-Point Adders
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
Optimizing Structured-Sparse Matrix Multiplication in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
GCN-ABFT: Low-Cost Online Error Checking for Graph Convolutional Networks
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Register Dispersion: Reducing the Footprint of the Vector Register File in Vector Engines of Low-Cost RISC-V CPUs
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Periodic Online Testing for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2025)
by: Peltekis, Christodoulos, et al.
Published: (2025)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
High-Performance Pipelined NTT Accelerators with Homogeneous Digit-Serial Modulo Arithmetic
by: Alexakis, George, et al.
Published: (2025)
by: Alexakis, George, et al.
Published: (2025)
Generalized Methodology for Determining Numerical Features of Hardware Floating-Point Matrix Multipliers: Part I
by: Khattak, Faizan A, et al.
Published: (2025)
by: Khattak, Faizan A, et al.
Published: (2025)
Efficient Implementation of RISC-V Vector Permutation Instructions
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines
by: Wang, Jiayi, et al.
Published: (2026)
by: Wang, Jiayi, et al.
Published: (2026)
TimeFloats: Train-in-Memory with Time-Domain Floating-Point Scalar Products
by: Hashem, Maeesha Binte, et al.
Published: (2024)
by: Hashem, Maeesha Binte, et al.
Published: (2024)
E2AFS: Energy-Efficient Approximate Floating Point Square Rooter for Error Tolerant Computing
by: Goyal, Prateek, et al.
Published: (2026)
by: Goyal, Prateek, et al.
Published: (2026)
A Stochastic Rounding-Enabled Low-Precision Floating-Point MAC for DNN Training
by: Ali, Sami Ben, et al.
Published: (2024)
by: Ali, Sami Ben, et al.
Published: (2024)
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
An Energy-Efficient Approximate Posit Multiply-Divide Unit
by: Thotli, Rishi, et al.
Published: (2026)
by: Thotli, Rishi, et al.
Published: (2026)
Low Power Approximate Multiplier Architecture for Deep Neural Networks
by: Jaswal, Pragun, et al.
Published: (2025)
by: Jaswal, Pragun, et al.
Published: (2025)
On Approximate 8-bit Floating-Point Operations Using Integer Operations
by: Lindberg, Theodor, et al.
Published: (2024)
by: Lindberg, Theodor, et al.
Published: (2024)
Fast Generation of Custom Floating-Point Spatial Filters on FPGAs
by: Campos, Nelson, et al.
Published: (2024)
by: Campos, Nelson, et al.
Published: (2024)
Floating Point HUB Adder for RISC-V Sargantana Processor
by: Bandera, Gerardo, et al.
Published: (2024)
by: Bandera, Gerardo, et al.
Published: (2024)
Hardware-Efficient CNNs: Interleaved Approximate FP32 Multipliers for Kernel Computation
by: Gowda, Bindu G, et al.
Published: (2025)
by: Gowda, Bindu G, et al.
Published: (2025)
Exploration of Unary Arithmetic-Based Matrix Multiply Units for Low Precision DL Accelerators
by: Vellaisamy, Prabhu, et al.
Published: (2026)
by: Vellaisamy, Prabhu, et al.
Published: (2026)
HPR-Mul: An Area and Energy-Efficient High-Precision Redundancy Multiplier by Approximate Computing
by: Vafaei, Jafar, et al.
Published: (2024)
by: Vafaei, Jafar, et al.
Published: (2024)
FPGA-Based Multiplier with a New Approximate Full Adder for Error-Resilient Applications
by: Ranjbar, Ali, et al.
Published: (2025)
by: Ranjbar, Ali, et al.
Published: (2025)
Small Logic-based Multipliers with Incomplete Sub-Multipliers for FPGAs
by: Böttcher, Andreas, et al.
Published: (2024)
by: Böttcher, Andreas, et al.
Published: (2024)
Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review
by: Gareau, Jaël Champagne, et al.
Published: (2026)
by: Gareau, Jaël Champagne, et al.
Published: (2026)
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
by: Natesh, Vikas, et al.
Published: (2025)
by: Natesh, Vikas, et al.
Published: (2025)
tubGEMM: Energy-Efficient and Sparsity-Effective Temporal-Unary-Binary Based Matrix Multiply Unit
by: Vellaisamy, Prabhu, et al.
Published: (2024)
by: Vellaisamy, Prabhu, et al.
Published: (2024)
Leveraging Highly Approximated Multipliers in DNN Inference
by: Zervakis, Georgios, et al.
Published: (2024)
by: Zervakis, Georgios, et al.
Published: (2024)
Revealing Untapped DSP Optimization Potentials for FPGA-Based Systolic Matrix Engines
by: Li, Jindong, et al.
Published: (2024)
by: Li, Jindong, et al.
Published: (2024)
BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
by: Han, Xiaomeng, et al.
Published: (2025)
by: Han, Xiaomeng, et al.
Published: (2025)
MXFormer: A Microscaling Floating-Point Charge-Trap Transistor Compute-in-Memory Transformer Accelerator
by: Karfakis, George, et al.
Published: (2026)
by: Karfakis, George, et al.
Published: (2026)
UFO-MAC: A Unified Framework for Optimization of High-Performance Multipliers and Multiply-Accumulators
by: Zuo, Dongsheng, et al.
Published: (2024)
by: Zuo, Dongsheng, et al.
Published: (2024)
A Bespoke Design Approach to Low-Power Printed Microprocessors for Machine Learning Applications
by: Chaidos, Panagiotis, et al.
Published: (2025)
by: Chaidos, Panagiotis, et al.
Published: (2025)
A Logic-Reuse Approach to Nibble-based Multiplier Design for Low Power Vector Computing
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026)
by: Chowdhury, Md Rownak Hossain, et al.
Published: (2026)
Similar Items
-
Error Checking for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2024) -
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024) -
Online Alignment and Addition in Multi-Term Floating-Point Adders
by: Alexandridis, Kosmas, et al.
Published: (2024) -
Optimizing Structured-Sparse Matrix Multiplication in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025) -
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)