Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Haiquan, Yao, Quanming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Graph Unitary Message Passing
by: Qiu, Haiquan, et al.
Published: (2024)
by: Qiu, Haiquan, et al.
Published: (2024)
Understanding Expressivity of GNN in Rule Learning
by: Qiu, Haiquan, et al.
Published: (2023)
by: Qiu, Haiquan, et al.
Published: (2023)
Superpose Task-specific Features for Model Merging
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026)
by: Chen, Yihong, et al.
Published: (2026)
Enhancing Training Efficiency Using Packing with Flash Attention
by: Kundu, Achintya, et al.
Published: (2024)
by: Kundu, Achintya, et al.
Published: (2024)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
by: Ke, Yekun, et al.
Published: (2024)
by: Ke, Yekun, et al.
Published: (2024)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024)
by: Shah, Jay, et al.
Published: (2024)
Why Do Transformers Fail to Forecast Time Series In-Context?
by: Zhou, Yufa, et al.
Published: (2025)
by: Zhou, Yufa, et al.
Published: (2025)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Accurate and interpretable drug-drug interaction prediction enabled by knowledge subgraph learning
by: Wang, Yaqing, et al.
Published: (2023)
by: Wang, Yaqing, et al.
Published: (2023)
PACIA: Parameter-Efficient Adapter for Few-Shot Molecular Property Prediction
by: Wu, Shiguang, et al.
Published: (2023)
by: Wu, Shiguang, et al.
Published: (2023)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time Series
by: Yu, Guoqi, et al.
Published: (2026)
by: Yu, Guoqi, et al.
Published: (2026)
Flash Invariant Point Attention
by: Liu, Andrew, et al.
Published: (2025)
by: Liu, Andrew, et al.
Published: (2025)
Neural Symbolic Regression of Complex Network Dynamics
by: Qiu, Haiquan, et al.
Published: (2024)
by: Qiu, Haiquan, et al.
Published: (2024)
Why Federated Optimization Fails to Achieve Perfect Fitting? A Theoretical Perspective on Client-Side Optima
by: Lei, Zhongxiang, et al.
Published: (2025)
by: Lei, Zhongxiang, et al.
Published: (2025)
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025)
by: Liao, Qichen, et al.
Published: (2025)
Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
FlashOptim: Optimizers for Memory-Efficient Training
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
Case-Based Reasoning Enhances the Predictive Power of LLMs in Drug-Drug Interaction
by: Liu, Guangyi, et al.
Published: (2025)
by: Liu, Guangyi, et al.
Published: (2025)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
by: Denisov-Blanch, Yegor, et al.
Published: (2026)
by: Denisov-Blanch, Yegor, et al.
Published: (2026)
Flash STU: Fast Spectral Transform Units
by: Liu, Y. Isabel, et al.
Published: (2024)
by: Liu, Y. Isabel, et al.
Published: (2024)
Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?
by: Yun, Vincent-Daniel
Published: (2025)
by: Yun, Vincent-Daniel
Published: (2025)
Scaling GraphLLM with Bilevel-Optimized Sparse Querying
by: Peng, Yangzhe, et al.
Published: (2026)
by: Peng, Yangzhe, et al.
Published: (2026)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
by: Sharma, Agniv, et al.
Published: (2024)
by: Sharma, Agniv, et al.
Published: (2024)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
by: Ding, Yifu, et al.
Published: (2026)
by: Ding, Yifu, et al.
Published: (2026)
VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation
by: Sun, Yupeng, et al.
Published: (2026)
by: Sun, Yupeng, et al.
Published: (2026)
Attending on Multilevel Structure of Proteins enables Accurate Prediction of Cold-Start Drug-Target Interactions
by: Zhang, Ziying, et al.
Published: (2025)
by: Zhang, Ziying, et al.
Published: (2025)
Automated Machine Learning: From Principles to Practices
by: Shen, Zhenqian, et al.
Published: (2018)
by: Shen, Zhenqian, et al.
Published: (2018)
Generalizing Hyperedge Expansion for Hyper-relational Knowledge Graph Modeling
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
A Versatile Graph Learning Approach through LLM-based Agent
by: Wei, Lanning, et al.
Published: (2023)
by: Wei, Lanning, et al.
Published: (2023)
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
Why Rectified Power Unit Networks Fail and How to Improve It: An Effective Field Theory Perspective
by: Kim, Taeyoung, et al.
Published: (2024)
by: Kim, Taeyoung, et al.
Published: (2024)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
by: Lu, Han, et al.
Published: (2025)
by: Lu, Han, et al.
Published: (2025)
FLASH-D: FlashAttention with Hidden Softmax Division
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
by: Gupta, Isha, et al.
Published: (2025)
by: Gupta, Isha, et al.
Published: (2025)
Similar Items
-
Graph Unitary Message Passing
by: Qiu, Haiquan, et al.
Published: (2024) -
Understanding Expressivity of GNN in Rule Learning
by: Qiu, Haiquan, et al.
Published: (2023) -
Superpose Task-specific Features for Model Merging
by: Qiu, Haiquan, et al.
Published: (2025) -
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
by: Chen, Yihong, et al.
Published: (2026) -
Enhancing Training Efficiency Using Packing with Flash Attention
by: Kundu, Achintya, et al.
Published: (2024)