AMLA: MUL by ADD in FlashAttention Rescaling
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Qichen, Hu, Chengqiu, Miao, Fangzheng, Li, Bao, Liu, Yiyang, Lyu, Junlong, Jiang, Lirui, Wang, Jun, Zheng, Lingchao, Li, Jun, Fan, Yuwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference
by: Zheng, Lingchao, et al.
Published: (2026)
by: Zheng, Lingchao, et al.
Published: (2026)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024)
by: Wang, Guoxia, et al.
Published: (2024)
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
by: Lin, Haoran, et al.
Published: (2024)
by: Lin, Haoran, et al.
Published: (2024)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Representation Shift: Unifying Token Compression with FlashAttention
by: Choi, Joonmyung, et al.
Published: (2025)
by: Choi, Joonmyung, et al.
Published: (2025)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024)
by: Shah, Jay, et al.
Published: (2024)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
by: Titopoulos, Vasileios, et al.
Published: (2025)
by: Titopoulos, Vasileios, et al.
Published: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
FlashAttention on a Napkin: A Diagrammatic Approach to Deep Learning IO-Awareness
by: Abbott, Vincent, et al.
Published: (2024)
by: Abbott, Vincent, et al.
Published: (2024)
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
by: Zadouri, Ted, et al.
Published: (2026)
by: Zadouri, Ted, et al.
Published: (2026)
LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model
by: Shao, Wei, et al.
Published: (2025)
by: Shao, Wei, et al.
Published: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
by: Zhou, Zhongchun, et al.
Published: (2026)
by: Zhou, Zhongchun, et al.
Published: (2026)
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
by: Font, Martí Llopart, et al.
Published: (2026)
by: Font, Martí Llopart, et al.
Published: (2026)
AIS: Adaptive Importance Sampling for Quantized RL
by: Zhou, Jiajun, et al.
Published: (2026)
by: Zhou, Jiajun, et al.
Published: (2026)
Online Pseudo-average Shifting Attention(PASA) for Robust Low-precision LLM Inference: Algorithms and Numerical Analysis
by: Cheng, Long, et al.
Published: (2025)
by: Cheng, Long, et al.
Published: (2025)
ADD-FWI
by: WU, YUPING
Published: (2025)
by: WU, YUPING
Published: (2025)
FlashGMM: Fast Gaussian Mixture Entropy Model for Learned Image Compression
by: Murai, Shimon, et al.
Published: (2025)
by: Murai, Shimon, et al.
Published: (2025)
CD1a affects the recurrence and prognosis of ovarian cancer
by: Qiong Zhu, et al.
Published: (2024)
by: Qiong Zhu, et al.
Published: (2024)
High-Layer Attention Pruning with Rescaling
by: Liu, Songtao, et al.
Published: (2025)
by: Liu, Songtao, et al.
Published: (2025)
The J=ÃDD¤ Vertex
by: F. S. Navarra
Published: (2004)
by: F. S. Navarra
Published: (2004)
Analyzing thermal‐moisture comfort and thermal protective performance of phase change materials dripped protective clothing
by: Zihan Gu, et al.
Published: (2024)
by: Zihan Gu, et al.
Published: (2024)
CircPTK2 as a Valuable Biomarker and Treatment Target in Cancer
by: Chengqiu Yan, et al.
Published: (2025)
by: Chengqiu Yan, et al.
Published: (2025)
A minimum problem associated with scalar Ginzburg-Landau equation and free boundary
by: Hu, Yuwei, et al.
Published: (2025)
by: Hu, Yuwei, et al.
Published: (2025)
Unpacking the Role of Embodied Pedagogical Agents in Multimedia Learning: Beyond Attention Guidance
by: Wenjing Li, et al.
Published: (2026)
by: Wenjing Li, et al.
Published: (2026)
RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
by: Fu, Ruibo, et al.
Published: (2025)
by: Fu, Ruibo, et al.
Published: (2025)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
by: Choi, Tae Eun, et al.
Published: (2026)
by: Choi, Tae Eun, et al.
Published: (2026)
ADD 2022: the First Audio Deep Synthesis Detection Challenge
by: Yi, Jiangyan, et al.
Published: (2022)
by: Yi, Jiangyan, et al.
Published: (2022)
Attention Sinks in Diffusion Transformers: A Causal Analysis
by: Wu, Fangzheng, et al.
Published: (2026)
by: Wu, Fangzheng, et al.
Published: (2026)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
ViT-MUL: A Baseline Study on Recent Machine Unlearning Methods Applied to Vision Transformers
by: Cho, Ikhyun, et al.
Published: (2024)
by: Cho, Ikhyun, et al.
Published: (2024)
MSV-Mamba: A Multiscale Vision Mamba Network for Echocardiography Segmentation
by: Yang, Xiaoxian, et al.
Published: (2025)
by: Yang, Xiaoxian, et al.
Published: (2025)
ADD for Multi-Bit Image Watermarking
by: Luo, An, et al.
Published: (2026)
by: Luo, An, et al.
Published: (2026)
Boundary-aware Decoupled Flow Networks for Realistic Extreme Rescaling
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
T-ADD: Enhancing DOA Estimation Robustness Against Adversarial Attacks
by: Zheng, Shilian, et al.
Published: (2025)
by: Zheng, Shilian, et al.
Published: (2025)
Is Flash Attention Stable?
by: Golden, Alicia, et al.
Published: (2024)
by: Golden, Alicia, et al.
Published: (2024)
Similar Items
-
Multi-Scale Dequant: Eliminating Dequantization Bottleneck via Activation Decomposition for Efficient LLM Inference
by: Zheng, Lingchao, et al.
Published: (2026) -
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024) -
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024) -
FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
by: Lin, Haoran, et al.
Published: (2024) -
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)