Flash STU: Fast Spectral Transform Units
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Y. Isabel, Nguyen, Windsor, Devre, Yagiz, Dogariu, Evan, Majumdar, Anirudha, Hazan, Elad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provable Length Generalization in Sequence Prediction via Spectral Filtering
by: Marsden, Annie, et al.
Published: (2024)
by: Marsden, Annie, et al.
Published: (2024)
Universal Learning of Nonlinear Dynamics
by: Dogariu, Evan, et al.
Published: (2025)
by: Dogariu, Evan, et al.
Published: (2025)
FutureFill: Fast Generation from Convolutional Sequence Models
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
by: Majumdar, Anirudha
Published: (2025)
by: Majumdar, Anirudha
Published: (2025)
SFO: Learning PDE Operators via Spectral Filtering
by: Koren, Noam, et al.
Published: (2026)
by: Koren, Noam, et al.
Published: (2026)
Research Program: Theory of Learning in Dynamical Systems
by: Hazan, Elad, et al.
Published: (2025)
by: Hazan, Elad, et al.
Published: (2025)
AI Alignment via Incentives and Correction
by: Agarwal, Rohit, et al.
Published: (2026)
by: Agarwal, Rohit, et al.
Published: (2026)
Thinking Forward and Backward: Effective Backward Planning with Large Language Models
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
The Hidden Game Problem
by: Buzaglo, Gon, et al.
Published: (2025)
by: Buzaglo, Gon, et al.
Published: (2025)
Spectral Filtering for Complex Linear Dynamical Systems
by: Hazan, Elad, et al.
Published: (2026)
by: Hazan, Elad, et al.
Published: (2026)
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast
by: Wu, Wenhao, et al.
Published: (2026)
by: Wu, Wenhao, et al.
Published: (2026)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024)
by: Shah, Jay, et al.
Published: (2024)
On The Statistical Representation Properties Of The Perturb-Softmax And The Perturb-Argmax Probability Distributions
by: Indelman, Hedda Cohen, et al.
Published: (2024)
by: Indelman, Hedda Cohen, et al.
Published: (2024)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
FlashSampling: Fast and Memory-Efficient Exact Sampling
by: Ruiz, Tomas, et al.
Published: (2026)
by: Ruiz, Tomas, et al.
Published: (2026)
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
by: Cohen, Ohad, et al.
Published: (2024)
by: Cohen, Ohad, et al.
Published: (2024)
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Predictive Red Teaming: Breaking Policies Without Breaking Robots
by: Majumdar, Anirudha, et al.
Published: (2025)
by: Majumdar, Anirudha, et al.
Published: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
The Spectral Lifecycle of Transformer Training: Transient Compression Waves, Persistent Spectral Gradients, and the Q/K--V Asymmetry
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
Flash Invariant Point Attention
by: Liu, Andrew, et al.
Published: (2025)
by: Liu, Andrew, et al.
Published: (2025)
Fast Graph Generation via Spectral Diffusion
by: Luo, Tianze, et al.
Published: (2022)
by: Luo, Tianze, et al.
Published: (2022)
Spectral State Space Models
by: Agarwal, Naman, et al.
Published: (2023)
by: Agarwal, Naman, et al.
Published: (2023)
AMLA: MUL by ADD in FlashAttention Rescaling
by: Liao, Qichen, et al.
Published: (2025)
by: Liao, Qichen, et al.
Published: (2025)
Spectral Transformer Neural Processes
by: Chen, Xianhe, et al.
Published: (2026)
by: Chen, Xianhe, et al.
Published: (2026)
The Phasor Transformer: Resolving Attention Bottlenecks on the Unit Circle
by: Sigdel, Dibakar
Published: (2026)
by: Sigdel, Dibakar
Published: (2026)
FlashOptim: Optimizers for Memory-Efficient Training
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
Universal Sequence Preconditioning
by: Marsden, Annie, et al.
Published: (2025)
by: Marsden, Annie, et al.
Published: (2025)
The Power of Second Order Methods for Sequence Preconditioning
by: Marsden, Annie, et al.
Published: (2026)
by: Marsden, Annie, et al.
Published: (2026)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
by: Lu, Han, et al.
Published: (2025)
by: Lu, Han, et al.
Published: (2025)
Spectral Text Fusion: A Frequency-Aware Approach to Multimodal Time-Series Forecasting
by: Nguyen, Huu Hiep, et al.
Published: (2026)
by: Nguyen, Huu Hiep, et al.
Published: (2026)
A Causal Framework to Measure and Mitigate Non-binary Treatment Discrimination
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
Enhancing Training Efficiency Using Packing with Flash Attention
by: Kundu, Achintya, et al.
Published: (2024)
by: Kundu, Achintya, et al.
Published: (2024)
Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence
by: Nguyen, Tien-Phat, et al.
Published: (2026)
by: Nguyen, Tien-Phat, et al.
Published: (2026)
Guiding Data Collection via Factored Scaling Curves
by: Zha, Lihan, et al.
Published: (2025)
by: Zha, Lihan, et al.
Published: (2025)
Explore until Confident: Efficient Exploration for Embodied Question Answering
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
Towards Flash Thinking via Decoupled Advantage Policy Optimization
by: Tan, Zezhong, et al.
Published: (2025)
by: Tan, Zezhong, et al.
Published: (2025)
Spectral Discovery of Continuous Symmetries via Generalized Fourier Transforms
by: Karjol, Pavan, et al.
Published: (2026)
by: Karjol, Pavan, et al.
Published: (2026)
How Transformers Learn In-Context Recall Tasks? Optimality, Training Dynamics and Generalization
by: Nguyen, Quan, et al.
Published: (2025)
by: Nguyen, Quan, et al.
Published: (2025)
From Universal to Individualized Actionability: Revisiting Personalization in Algorithmic Recourse
by: Budde, Lena Marie, et al.
Published: (2026)
by: Budde, Lena Marie, et al.
Published: (2026)
Similar Items
-
Provable Length Generalization in Sequence Prediction via Spectral Filtering
by: Marsden, Annie, et al.
Published: (2024) -
Universal Learning of Nonlinear Dynamics
by: Dogariu, Evan, et al.
Published: (2025) -
FutureFill: Fast Generation from Convolutional Sequence Models
by: Agarwal, Naman, et al.
Published: (2024) -
Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
by: Majumdar, Anirudha
Published: (2025) -
SFO: Learning PDE Operators via Spectral Filtering
by: Koren, Noam, et al.
Published: (2026)