Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
Fuente:
arXiv
Saved in:
| Main Authors: | Alexandridis, Kosmas, Titopoulos, Vasileios, Dimitrakopoulos, Giorgos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Register Dispersion: Reducing the Footprint of the Vector Register File in Vector Engines of Low-Cost RISC-V CPUs
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Optimizing Structured-Sparse Matrix Multiplication in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Online Alignment and Addition in Multi-Term Floating-Point Adders
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
GCN-ABFT: Low-Cost Online Error Checking for Graph Convolutional Networks
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Efficient Implementation of RISC-V Vector Permutation Instructions
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Periodic Online Testing for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2025)
by: Peltekis, Christodoulos, et al.
Published: (2025)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
by: Zhou, Zhongchun, et al.
Published: (2026)
by: Zhou, Zhongchun, et al.
Published: (2026)
MaRVIn: A Cross-Layer Mixed-Precision RISC-V Framework for DNN Inference, from ISA Extension to Hardware Acceleration
by: Armeniakos, Giorgos, et al.
Published: (2025)
by: Armeniakos, Giorgos, et al.
Published: (2025)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
by: Leon, Vasileios, et al.
Published: (2025)
by: Leon, Vasileios, et al.
Published: (2025)
High-Performance Pipelined NTT Accelerators with Homogeneous Digit-Serial Modulo Arithmetic
by: Alexakis, George, et al.
Published: (2025)
by: Alexakis, George, et al.
Published: (2025)
Error Checking for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
VeriBug: An Attention-based Framework for Bug-Localization in Hardware Designs
by: Stracquadanio, Giuseppe, et al.
Published: (2024)
by: Stracquadanio, Giuseppe, et al.
Published: (2024)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
by: Hsiung, Ching-Lin, et al.
Published: (2025)
by: Hsiung, Ching-Lin, et al.
Published: (2025)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
by: Font, Martí Llopart, et al.
Published: (2026)
by: Font, Martí Llopart, et al.
Published: (2026)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
by: Wang, Run, et al.
Published: (2025)
by: Wang, Run, et al.
Published: (2025)
Efficient and Reliable Vector Similarity Search Using Asymmetric Encoding with NAND-Flash for Many-Class Few-Shot Learning
by: Chiang, Hao-Wei, et al.
Published: (2024)
by: Chiang, Hao-Wei, et al.
Published: (2024)
A Joint Learning Approach to Hardware Caching and Prefetching
by: Yuan, Samuel, et al.
Published: (2025)
by: Yuan, Samuel, et al.
Published: (2025)
Energy-Aware Deep Learning on Resource-Constrained Hardware
by: Millar, Josh, et al.
Published: (2025)
by: Millar, Josh, et al.
Published: (2025)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
by: AbouElhamayed, Ahmed F., et al.
Published: (2023)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
by: Biggs, Benjamin, et al.
Published: (2023)
by: Biggs, Benjamin, et al.
Published: (2023)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
by: Yu, Zhewen, et al.
Published: (2024)
by: Yu, Zhewen, et al.
Published: (2024)
Taming the Exponential: A Fast Softmax Surrogate for Integer-Native Edge Inference
by: Danopoulos, Dimitrios, et al.
Published: (2026)
by: Danopoulos, Dimitrios, et al.
Published: (2026)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
by: Wang, Irene, et al.
Published: (2025)
by: Wang, Irene, et al.
Published: (2025)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
Hardware-Friendly Delayed-Feedback Reservoir for Multivariate Time-Series Classification
by: Ikeda, Sosei, et al.
Published: (2025)
by: Ikeda, Sosei, et al.
Published: (2025)
Hardware implementation of timely reliable Bayesian decision-making using memristors
by: Song, Lekai, et al.
Published: (2024)
by: Song, Lekai, et al.
Published: (2024)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
by: Zhang, Zehuan, et al.
Published: (2024)
by: Zhang, Zehuan, et al.
Published: (2024)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
by: Zhang, Zehuan, et al.
Published: (2026)
by: Zhang, Zehuan, et al.
Published: (2026)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
by: Choi, Dawon, et al.
Published: (2026)
by: Choi, Dawon, et al.
Published: (2026)
Similar Items
-
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025) -
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025) -
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025) -
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025) -
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)