FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Zadouri, Ted, Hoehnerbach, Markus, Shah, Jay, Liu, Timmy, Thakkar, Vijay, Dao, Tri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hardware-Efficient Attention for Fast Decoding
by: Zadouri, Ted, et al.
Published: (2025)
by: Zadouri, Ted, et al.
Published: (2025)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024)
by: Shah, Jay, et al.
Published: (2024)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025)
by: Liao, Qichen, et al.
Published: (2025)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024)
by: Wang, Guoxia, et al.
Published: (2024)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
by: Zhou, Zhongchun, et al.
Published: (2026)
by: Zhou, Zhongchun, et al.
Published: (2026)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Representation Shift: Unifying Token Compression with FlashAttention
by: Choi, Joonmyung, et al.
Published: (2025)
by: Choi, Joonmyung, et al.
Published: (2025)
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
by: Lin, Haoran, et al.
Published: (2024)
by: Lin, Haoran, et al.
Published: (2024)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
by: Yuan, Jiayi, et al.
Published: (2025)
by: Yuan, Jiayi, et al.
Published: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
FlashAttention on a Napkin: A Diagrammatic Approach to Deep Learning IO-Awareness
by: Abbott, Vincent, et al.
Published: (2024)
by: Abbott, Vincent, et al.
Published: (2024)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
by: Ding, Yifu, et al.
Published: (2026)
by: Ding, Yifu, et al.
Published: (2026)
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
by: Sun, Luoyang, et al.
Published: (2026)
by: Sun, Luoyang, et al.
Published: (2026)
Abstractions-of-Thought: Intermediate Representations for LLM Reasoning in Hardware Design
by: DeLorenzo, Matthew, et al.
Published: (2025)
by: DeLorenzo, Matthew, et al.
Published: (2025)
AdaSplash: Adaptive Sparse Flash Attention
by: Gonçalves, Nuno, et al.
Published: (2025)
by: Gonçalves, Nuno, et al.
Published: (2025)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025)
by: Nrusimha, Aniruddha, et al.
Published: (2025)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
by: Sharma, Agniv, et al.
Published: (2024)
by: Sharma, Agniv, et al.
Published: (2024)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
by: Font, Martí Llopart, et al.
Published: (2026)
by: Font, Martí Llopart, et al.
Published: (2026)
FlashEVA: Accelerating LLM inference via Efficient Attention
by: Kostelec, Juan Gabriel, et al.
Published: (2025)
by: Kostelec, Juan Gabriel, et al.
Published: (2025)
WISTERIA: Weak Implicit Signal-based Temporal Relation Extraction with Attention
by: Do, Duy Dao, et al.
Published: (2026)
by: Do, Duy Dao, et al.
Published: (2026)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
by: Lin, Xihui, et al.
Published: (2024)
by: Lin, Xihui, et al.
Published: (2024)
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
by: Deng, Cheng, et al.
Published: (2025)
by: Deng, Cheng, et al.
Published: (2025)
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency
by: Cai, Aichen, et al.
Published: (2026)
by: Cai, Aichen, et al.
Published: (2026)
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
by: Guo, Han, et al.
Published: (2026)
by: Guo, Han, et al.
Published: (2026)
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
by: Paliotta, Daniele, et al.
Published: (2025)
by: Paliotta, Daniele, et al.
Published: (2025)
Arch: An AI-Native Hardware Description Language for Register-Transfer Clocked Hardware Design
by: Zhao, Shuqing
Published: (2026)
by: Zhao, Shuqing
Published: (2026)
Attention Is All You Need for KV Cache in Diffusion LLMs
by: Nguyen-Tri, Quan, et al.
Published: (2025)
by: Nguyen-Tri, Quan, et al.
Published: (2025)
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
by: Dao, Tri, et al.
Published: (2024)
by: Dao, Tri, et al.
Published: (2024)
DFlash: Block Diffusion for Flash Speculative Decoding
by: Chen, Jian, et al.
Published: (2026)
by: Chen, Jian, et al.
Published: (2026)
Evaluating LLMs for Hardware Design and Test
by: Blocklove, Jason, et al.
Published: (2024)
by: Blocklove, Jason, et al.
Published: (2024)
Bhasha-Rupantarika: Algorithm-Hardware Co-design approach for Multilingual Neural Machine Translation
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
SemEval-2017 Task 4: Sentiment Analysis in Twitter using BERT
by: Das, Rupak Kumar, et al.
Published: (2024)
by: Das, Rupak Kumar, et al.
Published: (2024)
Similar Items
-
Hardware-Efficient Attention for Fast Decoding
by: Zadouri, Ted, et al.
Published: (2025) -
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024) -
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025) -
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024) -
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025)