FLASH-D: FlashAttention with Hidden Softmax Division
Fuente:
arXiv
Saved in:
| Main Authors: | Alexandridis, Kosmas, Titopoulos, Vasileios, Dimitrakopoulos, Giorgos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Register Dispersion: Reducing the Footprint of the Vector Register File in Vector Engines of Low-Cost RISC-V CPUs
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Optimizing Structured-Sparse Matrix Multiplication in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Online Alignment and Addition in Multi-Term Floating-Point Adders
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
Floating-Point Multiply-Add with Approximate Normalization for Low-Cost Matrix Engines
by: Alexandridis, Kosmas, et al.
Published: (2024)
by: Alexandridis, Kosmas, et al.
Published: (2024)
Efficient Implementation of RISC-V Vector Permutation Instructions
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
GCN-ABFT: Low-Cost Online Error Checking for Graph Convolutional Networks
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Periodic Online Testing for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2025)
by: Peltekis, Christodoulos, et al.
Published: (2025)
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
by: Liu, Shiwei, et al.
Published: (2024)
by: Liu, Shiwei, et al.
Published: (2024)
Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
by: Huang, Zhen, et al.
Published: (2025)
by: Huang, Zhen, et al.
Published: (2025)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
by: İslamoğlu, Gamze, et al.
Published: (2023)
by: İslamoğlu, Gamze, et al.
Published: (2023)
Dynamic Sparse Attention: Access Patterns and Architecture
by: Levy, Noam
Published: (2026)
by: Levy, Noam
Published: (2026)
In-Memory Learning Automata Architecture using Y-Flash Cell
by: Ghazal, Omar, et al.
Published: (2024)
by: Ghazal, Omar, et al.
Published: (2024)
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
by: Kabir, Ehsan, et al.
Published: (2024)
by: Kabir, Ehsan, et al.
Published: (2024)
Revolutionizing TCAD Simulations with Universal Device Encoding and Graph Attention Networks
by: Fan, Guangxi, et al.
Published: (2023)
by: Fan, Guangxi, et al.
Published: (2023)
Reducing the Cost of Dropout in Flash-Attention by Hiding RNG with GEMM
by: Ma, Haiyue, et al.
Published: (2024)
by: Ma, Haiyue, et al.
Published: (2024)
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
by: Zhou, Zhongchun, et al.
Published: (2026)
by: Zhou, Zhongchun, et al.
Published: (2026)
Multi-objective Optimization in CPU Design Space Exploration: Attention is All You Need
by: Xue, Runzhen, et al.
Published: (2024)
by: Xue, Runzhen, et al.
Published: (2024)
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
by: Cong, Rongqing, et al.
Published: (2024)
by: Cong, Rongqing, et al.
Published: (2024)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
High-Performance Pipelined NTT Accelerators with Homogeneous Digit-Serial Modulo Arithmetic
by: Alexakis, George, et al.
Published: (2025)
by: Alexakis, George, et al.
Published: (2025)
Error Checking for Sparse Systolic Tensor Arrays
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
by: Rakka, Mariam, et al.
Published: (2024)
by: Rakka, Mariam, et al.
Published: (2024)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
by: Choi, Dawon, et al.
Published: (2026)
by: Choi, Dawon, et al.
Published: (2026)
Optimization of bi-directional gated loop cell based on multi-head attention mechanism for SSD health state classification model
by: Wen, Zhizhao, et al.
Published: (2025)
by: Wen, Zhizhao, et al.
Published: (2025)
On-Sensor Convolutional Neural Networks with Early-Exits
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
by: Liang, Yanbiao, et al.
Published: (2025)
by: Liang, Yanbiao, et al.
Published: (2025)
Efficient Calibration for RRAM-based In-Memory Computing using DoRA
by: Dong, Weirong, et al.
Published: (2025)
by: Dong, Weirong, et al.
Published: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
Neural Network Acceleration on MPSoC board: Integrating SLAC's SNL, Rogue Software and Auto-SNL
by: Rahali, Hamza Ezzaoui, et al.
Published: (2025)
by: Rahali, Hamza Ezzaoui, et al.
Published: (2025)
GATMesh: Clock Mesh Timing Analysis using Graph Neural Networks
by: Khan, Muhammad Hadir, et al.
Published: (2025)
by: Khan, Muhammad Hadir, et al.
Published: (2025)
Extending Straight-Through Estimation for Robust Neural Networks on Analog CIM Hardware
by: Feng, Yuannuo, et al.
Published: (2025)
by: Feng, Yuannuo, et al.
Published: (2025)
Similar Items
-
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025) -
Custom Algorithm-based Fault Tolerance for Attention Layers in Transformers
by: Titopoulos, Vasileios, et al.
Published: (2025) -
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025) -
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024) -
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)