FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Fuente:
arXiv
Saved in:
| Main Authors: | Shah, Jay, Bikshandi, Ganesh, Zhang, Ying, Thakkar, Vijay, Ramani, Pradeep, Dao, Tri |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025)
by: Liao, Qichen, et al.
Published: (2025)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
by: Zadouri, Ted, et al.
Published: (2026)
by: Zadouri, Ted, et al.
Published: (2026)
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
by: Lin, Haoran, et al.
Published: (2024)
by: Lin, Haoran, et al.
Published: (2024)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
by: Lu, Han, et al.
Published: (2025)
by: Lu, Han, et al.
Published: (2025)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024)
by: Wang, Guoxia, et al.
Published: (2024)
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Flash Invariant Point Attention
by: Liu, Andrew, et al.
Published: (2025)
by: Liu, Andrew, et al.
Published: (2025)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
Hardware-Aware Reformulation of Convolutions for Efficient Execution on Specialized AI Hardware: A Case Study on NVIDIA Tensor Cores
by: Bikshandi, Ganesh
Published: (2026)
by: Bikshandi, Ganesh
Published: (2026)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
by: Sharma, Agniv, et al.
Published: (2024)
by: Sharma, Agniv, et al.
Published: (2024)
Enhancing Training Efficiency Using Packing with Flash Attention
by: Kundu, Achintya, et al.
Published: (2024)
by: Kundu, Achintya, et al.
Published: (2024)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
by: Gu, Albert, et al.
Published: (2023)
by: Gu, Albert, et al.
Published: (2023)
VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation
by: Sun, Yupeng, et al.
Published: (2026)
by: Sun, Yupeng, et al.
Published: (2026)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
FAST: Factorizable Attention for Speeding up Transformers
by: Gerami, Armin, et al.
Published: (2024)
by: Gerami, Armin, et al.
Published: (2024)
MAC-Attention: a Match-Amend-Complete Scheme for Fast and Accurate Attention Computation
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
Hardware-Efficient Attention for Fast Decoding
by: Zadouri, Ted, et al.
Published: (2025)
by: Zadouri, Ted, et al.
Published: (2025)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
FlashAttention on a Napkin: A Diagrammatic Approach to Deep Learning IO-Awareness
by: Abbott, Vincent, et al.
Published: (2024)
by: Abbott, Vincent, et al.
Published: (2024)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash Attention
by: Xu, Rengan, et al.
Published: (2024)
by: Xu, Rengan, et al.
Published: (2024)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
by: Huang, Luke J., et al.
Published: (2026)
by: Huang, Luke J., et al.
Published: (2026)
Attention Is All You Need for KV Cache in Diffusion LLMs
by: Nguyen-Tri, Quan, et al.
Published: (2025)
by: Nguyen-Tri, Quan, et al.
Published: (2025)
Multi-branch of Attention Yields Accurate Results for Tabular Data
by: Li, Xuechen, et al.
Published: (2025)
by: Li, Xuechen, et al.
Published: (2025)
Fast and Simplex: 2-Simplicial Attention in Triton
by: Roy, Aurko, et al.
Published: (2025)
by: Roy, Aurko, et al.
Published: (2025)
Flash STU: Fast Spectral Transform Units
by: Liu, Y. Isabel, et al.
Published: (2024)
by: Liu, Y. Isabel, et al.
Published: (2024)
Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning
by: Lu, Jian
Published: (2025)
by: Lu, Jian
Published: (2025)
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
by: Hassani, Ali, et al.
Published: (2025)
by: Hassani, Ali, et al.
Published: (2025)
SFC: Achieve Accurate Fast Convolution under Low-precision Arithmetic
by: He, Liulu, et al.
Published: (2024)
by: He, Liulu, et al.
Published: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers
by: Hwang, Sukjun, et al.
Published: (2024)
by: Hwang, Sukjun, et al.
Published: (2024)
SageBwd: A Trainable Low-bit Attention
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026)
by: Luo, Haozheng, et al.
Published: (2026)
Fast and Accurate Identification of Hardware Trojan Locations in Gate-Level Netlist using Nearest Neighbour Approach integrated with Machine Learning Technique
by: Chattopadhyay, Anindita, et al.
Published: (2025)
by: Chattopadhyay, Anindita, et al.
Published: (2025)
Similar Items
-
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024) -
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025) -
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025) -
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
by: Zadouri, Ted, et al.
Published: (2026) -
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
by: Lin, Haoran, et al.
Published: (2024)