FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Haoran, Yu, Xianzhi, Zhao, Kang, Hou, Lu, Zhan, Zongyuan, Kamenev, Stanislav, Bao, Han, Hu, Ting, Wang, Mingkai, Chang, Qixin, Sui, Siyue, Sun, Weihao, Hu, Jiaxin, Yao, Jun, Yin, Zekun, Qian, Cheng, Zhang, Ying, Pan, Yinfei, Yang, Yu, Liu, Weiguo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025)
by: Liao, Qichen, et al.
Published: (2025)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024)
by: Wang, Guoxia, et al.
Published: (2024)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024)
by: Shah, Jay, et al.
Published: (2024)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Representation Shift: Unifying Token Compression with FlashAttention
by: Choi, Joonmyung, et al.
Published: (2025)
by: Choi, Joonmyung, et al.
Published: (2025)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
by: Lin, Haoran, et al.
Published: (2025)
by: Lin, Haoran, et al.
Published: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
FlashAttention on a Napkin: A Diagrammatic Approach to Deep Learning IO-Awareness
by: Abbott, Vincent, et al.
Published: (2024)
by: Abbott, Vincent, et al.
Published: (2024)
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
by: Zadouri, Ted, et al.
Published: (2026)
by: Zadouri, Ted, et al.
Published: (2026)
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
by: Zhou, Zhongchun, et al.
Published: (2026)
by: Zhou, Zhongchun, et al.
Published: (2026)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
by: Font, Martí Llopart, et al.
Published: (2026)
by: Font, Martí Llopart, et al.
Published: (2026)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
by: Deshmukh, Aadesh, et al.
Published: (2025)
by: Deshmukh, Aadesh, et al.
Published: (2025)
DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
Is Flash Attention Stable?
by: Golden, Alicia, et al.
Published: (2024)
by: Golden, Alicia, et al.
Published: (2024)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
by: Dege, Pengcuo, et al.
Published: (2025)
by: Dege, Pengcuo, et al.
Published: (2025)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
by: Huang, Zongle, et al.
Published: (2025)
by: Huang, Zongle, et al.
Published: (2025)
Flash Invariant Point Attention
by: Liu, Andrew, et al.
Published: (2025)
by: Liu, Andrew, et al.
Published: (2025)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
The I/O Complexity of Attention, or How Optimal is Flash Attention?
by: Saha, Barna, et al.
Published: (2024)
by: Saha, Barna, et al.
Published: (2024)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
by: Sharma, Agniv, et al.
Published: (2024)
by: Sharma, Agniv, et al.
Published: (2024)
Perspectivas de la Innovación de Políticas Pública en China
by: Hu Xianzhi
Published: (2020)
by: Hu Xianzhi
Published: (2020)
FlashDecoding++: Faster Large Language Model Inference on GPUs
by: Hong, Ke, et al.
Published: (2023)
by: Hong, Ke, et al.
Published: (2023)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
by: Zhao, Pengxiang, et al.
Published: (2026)
by: Zhao, Pengxiang, et al.
Published: (2026)
AdaSplash: Adaptive Sparse Flash Attention
by: Gonçalves, Nuno, et al.
Published: (2025)
by: Gonçalves, Nuno, et al.
Published: (2025)
FlashBias: Fast Computation of Attention with Bias
by: Wu, Haixu, et al.
Published: (2025)
by: Wu, Haixu, et al.
Published: (2025)
E$^3$-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
by: Yuan, Tao, et al.
Published: (2025)
by: Yuan, Tao, et al.
Published: (2025)
Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers
by: Liu, Yuxi, et al.
Published: (2026)
by: Liu, Yuxi, et al.
Published: (2026)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
by: Jin, Zehao, et al.
Published: (2026)
by: Jin, Zehao, et al.
Published: (2026)
EAGLE-Pangu: Accelerator-Safe Tree Speculative Decoding on Ascend NPUs
by: Han, Chang, et al.
Published: (2026)
by: Han, Chang, et al.
Published: (2026)
Enhancing Training Efficiency Using Packing with Flash Attention
by: Kundu, Achintya, et al.
Published: (2024)
by: Kundu, Achintya, et al.
Published: (2024)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
by: Choudhary, Mansi, et al.
Published: (2025)
by: Choudhary, Mansi, et al.
Published: (2025)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
by: Fu, Tsu-Jui, et al.
Published: (2023)
by: Fu, Tsu-Jui, et al.
Published: (2023)
Similar Items
-
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025) -
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024) -
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024) -
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024) -
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)