Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Sobczyk, Aleksandros, Gottardo, Gioele, Matzoros, Christos K., De Vita, Mirko, Skogh, Filip, Zouzias, Anastasios, Zhuang, Jiawei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prefix Sums via Kronecker Products
by: Sobczyk, Aleksandros, et al.
Published: (2025)
by: Sobczyk, Aleksandros, et al.
Published: (2025)
Approaching I/O-optimality for Approximate Attention
by: Papp, Pál András, et al.
Published: (2026)
by: Papp, Pál András, et al.
Published: (2026)
Segmented Operations using Matrix Multiplications
by: Sobczyk, Aleksandros, et al.
Published: (2025)
by: Sobczyk, Aleksandros, et al.
Published: (2025)
Parallel Scan on Ascend AI Accelerators
by: Wróblewski, Bartłomiej, et al.
Published: (2025)
by: Wróblewski, Bartłomiej, et al.
Published: (2025)
Quantum Doubly Stochastic Transformers
by: Born, Jannis, et al.
Published: (2025)
by: Born, Jannis, et al.
Published: (2025)
I/O complexity and pebble games with partial computations
by: Sobczyk, Aleksandros
Published: (2024)
by: Sobczyk, Aleksandros
Published: (2024)
Deterministic complexity analysis of Hermitian eigenproblems
by: Sobczyk, Aleksandros
Published: (2024)
by: Sobczyk, Aleksandros
Published: (2024)
Spectral Gaps with Quantum Counting Queries and Oblivious State Preparation
by: Vazquez, Almudena Carrera, et al.
Published: (2025)
by: Vazquez, Almudena Carrera, et al.
Published: (2025)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
Invariant subspaces and PCA in nearly matrix multiplication time
by: Sobczyk, Aleksandros, et al.
Published: (2023)
by: Sobczyk, Aleksandros, et al.
Published: (2023)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
A Parallel Scan Algorithm in the Tensor Core Unit Model
by: Zouzias, Anastasios, et al.
Published: (2024)
by: Zouzias, Anastasios, et al.
Published: (2024)
The Impact of Partial Computations on the Red-Blue Pebble Game
by: Papp, Pál András, et al.
Published: (2025)
by: Papp, Pál András, et al.
Published: (2025)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
by: Zhou, Chenyu, et al.
Published: (2026)
by: Zhou, Chenyu, et al.
Published: (2026)
FlashNorm: Fast Normalization for Transformers
by: Graef, Nils, et al.
Published: (2024)
by: Graef, Nils, et al.
Published: (2024)
Gated Delta Networks: Improving Mamba2 with Delta Rule
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
Transportation mode recognition based on low-rate acceleration and location signals with an attention-based multiple-instance learning network
by: Siargkas, Christos, et al.
Published: (2024)
by: Siargkas, Christos, et al.
Published: (2024)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
by: Willette, Jeffrey, et al.
Published: (2025)
by: Willette, Jeffrey, et al.
Published: (2025)
RIS-Assisted 3D Spherical Splatting for Object Composition Visualization using Detection Transformers
by: Sotiropoulos, Anastasios T., et al.
Published: (2025)
by: Sotiropoulos, Anastasios T., et al.
Published: (2025)
On Maximum Entropy Linear Feature Inversion
by: Baggenstoss, Paul M
Published: (2024)
by: Baggenstoss, Paul M
Published: (2024)
FlashAttention on a Napkin: A Diagrammatic Approach to Deep Learning IO-Awareness
by: Abbott, Vincent, et al.
Published: (2024)
by: Abbott, Vincent, et al.
Published: (2024)
Deceptron: Learned Local Inverses for Fast and Stable Physics Inversion
by: Kachhadiya, Aaditya L.
Published: (2025)
by: Kachhadiya, Aaditya L.
Published: (2025)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
by: Mantes, Albert Dominguez, et al.
Published: (2026)
by: Mantes, Albert Dominguez, et al.
Published: (2026)
Weaves, Wires, and Morphisms: Formalizing and Implementing the Algebra of Deep Learning
by: Abbott, Vincent, et al.
Published: (2026)
by: Abbott, Vincent, et al.
Published: (2026)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
by: Pandey, Vishal, et al.
Published: (2026)
by: Pandey, Vishal, et al.
Published: (2026)
Diffusion Model Guided Sampling with Pixel-Wise Aleatoric Uncertainty Estimation
by: De Vita, Michele, et al.
Published: (2024)
by: De Vita, Michele, et al.
Published: (2024)
Sparse Hybrid Linear-Morphological Networks
by: Fotopoulos, Konstantinos, et al.
Published: (2025)
by: Fotopoulos, Konstantinos, et al.
Published: (2025)
Forecasting the Past: Gradient-Based Distribution Shift Detection in Trajectory Prediction
by: De Vita, Michele, et al.
Published: (2026)
by: De Vita, Michele, et al.
Published: (2026)
FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
Generative Inversion of Spectroscopic Data for Amorphous Structure Elucidation
by: Guo, Jiawei, et al.
Published: (2026)
by: Guo, Jiawei, et al.
Published: (2026)
Fast Linear Solvers via AI-Tuned Markov Chain Monte Carlo-based Matrix Inversion
by: Lebedev, Anton, et al.
Published: (2025)
by: Lebedev, Anton, et al.
Published: (2025)
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
by: Huang, Yulong, et al.
Published: (2026)
by: Huang, Yulong, et al.
Published: (2026)
TopoMap: A Feature-based Semantic Discriminator of the Topographical Regions in the Test Input Space
by: De Vita, Gianmarco, et al.
Published: (2025)
by: De Vita, Gianmarco, et al.
Published: (2025)
Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences
by: Tumma, Neehal, et al.
Published: (2026)
by: Tumma, Neehal, et al.
Published: (2026)
Fast Clustering of Categorical Big Data
by: Thapaliya, Bipana, et al.
Published: (2025)
by: Thapaliya, Bipana, et al.
Published: (2025)
ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers
by: Hsu, Chih-Chung, et al.
Published: (2026)
by: Hsu, Chih-Chung, et al.
Published: (2026)
Fast and Robust Simulation-Based Inference With Optimization Monte Carlo
by: Gkolemis, Vasilis, et al.
Published: (2025)
by: Gkolemis, Vasilis, et al.
Published: (2025)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Seamlessly Integrating Tree-Based Positional Embeddings into Transformer Models for Source Code Representation
by: Bartkowiak, Patryk, et al.
Published: (2025)
by: Bartkowiak, Patryk, et al.
Published: (2025)
Similar Items
-
Prefix Sums via Kronecker Products
by: Sobczyk, Aleksandros, et al.
Published: (2025) -
Approaching I/O-optimality for Approximate Attention
by: Papp, Pál András, et al.
Published: (2026) -
Segmented Operations using Matrix Multiplications
by: Sobczyk, Aleksandros, et al.
Published: (2025) -
Parallel Scan on Ascend AI Accelerators
by: Wróblewski, Bartłomiej, et al.
Published: (2025) -
Quantum Doubly Stochastic Transformers
by: Born, Jannis, et al.
Published: (2025)