Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
Fuente:
arXiv
Guardado en:
| Autores principales: | Beck, Maximilian, Pöppel, Korbinian, Lippe, Phillip, Hochreiter, Sepp |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware
por: Pöppel, Korbinian, et al.
Publicado: (2024)
por: Pöppel, Korbinian, et al.
Publicado: (2024)
Vision-LSTM: xLSTM as Generic Vision Backbone
por: Alkin, Benedikt, et al.
Publicado: (2024)
por: Alkin, Benedikt, et al.
Publicado: (2024)
xLSTM 7B: A Recurrent LLM for Fast and Efficient Inference
por: Beck, Maximilian, et al.
Publicado: (2025)
por: Beck, Maximilian, et al.
Publicado: (2025)
xLSTM: Extended Long Short-Term Memory
por: Beck, Maximilian, et al.
Publicado: (2024)
por: Beck, Maximilian, et al.
Publicado: (2024)
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
por: Schmied, Thomas, et al.
Publicado: (2024)
por: Schmied, Thomas, et al.
Publicado: (2024)
xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
por: Beck, Maximilian, et al.
Publicado: (2025)
por: Beck, Maximilian, et al.
Publicado: (2025)
pLSTM: parallelizable Linear Source Transition Mark networks
por: Pöppel, Korbinian, et al.
Publicado: (2025)
por: Pöppel, Korbinian, et al.
Publicado: (2025)
Bio-xLSTM: Generative modeling, representation and in-context learning of biological and chemical sequences
por: Schmidinger, Niklas, et al.
Publicado: (2024)
por: Schmidinger, Niklas, et al.
Publicado: (2024)
Effective Distillation to Hybrid xLSTM Architectures
por: Hauzenberger, Lukas, et al.
Publicado: (2026)
por: Hauzenberger, Lukas, et al.
Publicado: (2026)
xLSTM-ECG: Multi-label ECG Classification via Feature Fusion with xLSTM
por: Kang, Lei, et al.
Publicado: (2025)
por: Kang, Lei, et al.
Publicado: (2025)
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures
por: Thiombiano, Abdoul Majid O., et al.
Publicado: (2025)
por: Thiombiano, Abdoul Majid O., et al.
Publicado: (2025)
xLSTMTime : Long-term Time Series Forecasting With xLSTM
por: Alharthi, Musleh, et al.
Publicado: (2024)
por: Alharthi, Musleh, et al.
Publicado: (2024)
xLSTMAD: A Powerful xLSTM-based Method for Anomaly Detection
por: Faber, Kamil, et al.
Publicado: (2025)
por: Faber, Kamil, et al.
Publicado: (2025)
Enhancing Spatiotemporal Networks with xLSTM: A Scalar LSTM Approach for Cellular Traffic Forecasting
por: Ali, Khalid, et al.
Publicado: (2025)
por: Ali, Khalid, et al.
Publicado: (2025)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
por: Thiombiano, Abdoul Majid O., et al.
Publicado: (2025)
por: Thiombiano, Abdoul Majid O., et al.
Publicado: (2025)
Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images
por: Zhu, Qinfeng, et al.
Publicado: (2024)
por: Zhu, Qinfeng, et al.
Publicado: (2024)
TiledAttention: a CUDA Tile SDPA Kernel for PyTorch
por: Khan, Taimur
Publicado: (2026)
por: Khan, Taimur
Publicado: (2026)
Adaptive Retrieval helps Reasoning in LLMs -- but mostly if it's not used
por: Shakya, Srijan, et al.
Publicado: (2026)
por: Shakya, Srijan, et al.
Publicado: (2026)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
por: Kühne, Nikolai Lund, et al.
Publicado: (2025)
por: Kühne, Nikolai Lund, et al.
Publicado: (2025)
SKOLR: Structured Koopman Operator Linear RNN for Time-Series Forecasting
por: Zhang, Yitian, et al.
Publicado: (2025)
por: Zhang, Yitian, et al.
Publicado: (2025)
A Diffusion Model Framework for Unsupervised Neural Combinatorial Optimization
por: Sanokowski, Sebastian, et al.
Publicado: (2024)
por: Sanokowski, Sebastian, et al.
Publicado: (2024)
Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language Generation
por: Ielanskyi, Mykyta, et al.
Publicado: (2025)
por: Ielanskyi, Mykyta, et al.
Publicado: (2025)
On Information-Theoretic Measures of Predictive Uncertainty
por: Schweighofer, Kajetan, et al.
Publicado: (2024)
por: Schweighofer, Kajetan, et al.
Publicado: (2024)
Improving Uncertainty Estimation through Semantically Diverse Language Generation
por: Aichberger, Lukas, et al.
Publicado: (2024)
por: Aichberger, Lukas, et al.
Publicado: (2024)
Contrastive Abstraction for Reinforcement Learning
por: Patil, Vihang, et al.
Publicado: (2024)
por: Patil, Vihang, et al.
Publicado: (2024)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
por: You, Haoran, et al.
Publicado: (2024)
por: You, Haoran, et al.
Publicado: (2024)
The Disparate Benefits of Deep Ensembles
por: Schweighofer, Kajetan, et al.
Publicado: (2024)
por: Schweighofer, Kajetan, et al.
Publicado: (2024)
Linear Attention for Efficient Bidirectional Sequence Modeling
por: Afzal, Arshia, et al.
Publicado: (2025)
por: Afzal, Arshia, et al.
Publicado: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
por: Zhu, Yifan, et al.
Publicado: (2026)
por: Zhu, Yifan, et al.
Publicado: (2026)
SLAY: Geometry-Aware Spherical Linearized Attention with Yat-Kernel
por: Luna, Jose Miguel, et al.
Publicado: (2026)
por: Luna, Jose Miguel, et al.
Publicado: (2026)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
por: Mishra, Mayank, et al.
Publicado: (2026)
por: Mishra, Mayank, et al.
Publicado: (2026)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
por: Chen, Shimao, et al.
Publicado: (2024)
por: Chen, Shimao, et al.
Publicado: (2024)
Rethinking Losses for Diffusion Bridge Samplers
por: Sanokowski, Sebastian, et al.
Publicado: (2025)
por: Sanokowski, Sebastian, et al.
Publicado: (2025)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
por: Lu, Jiecheng, et al.
Publicado: (2026)
por: Lu, Jiecheng, et al.
Publicado: (2026)
Exact Linear Attention
por: Ou, Weinuo
Publicado: (2026)
por: Ou, Weinuo
Publicado: (2026)
Kaczmarz Linear Attention
por: Zou, Jiaxuan, et al.
Publicado: (2026)
por: Zou, Jiaxuan, et al.
Publicado: (2026)
A Comprehensive Overview and Comparative Analysis on Deep Learning Models: CNN, RNN, LSTM, GRU
por: Shiri, Farhad Mortezapour, et al.
Publicado: (2023)
por: Shiri, Farhad Mortezapour, et al.
Publicado: (2023)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
por: Zuo, Yifei, et al.
Publicado: (2025)
por: Zuo, Yifei, et al.
Publicado: (2025)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
por: Hu, Wenjie, et al.
Publicado: (2025)
por: Hu, Wenjie, et al.
Publicado: (2025)
Parameter Efficient Fine-tuning via Explained Variance Adaptation
por: Paischer, Fabian, et al.
Publicado: (2024)
por: Paischer, Fabian, et al.
Publicado: (2024)
Ejemplares similares
-
FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware
por: Pöppel, Korbinian, et al.
Publicado: (2024) -
Vision-LSTM: xLSTM as Generic Vision Backbone
por: Alkin, Benedikt, et al.
Publicado: (2024) -
xLSTM 7B: A Recurrent LLM for Fast and Efficient Inference
por: Beck, Maximilian, et al.
Publicado: (2025) -
xLSTM: Extended Long Short-Term Memory
por: Beck, Maximilian, et al.
Publicado: (2024) -
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
por: Schmied, Thomas, et al.
Publicado: (2024)