Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sobczyk, Aleksandros, Gottardo, Gioele, Matzoros, Christos K., De Vita, Mirko, Skogh, Filip, Zouzias, Anastasios, Zhuang, Jiawei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910241238548480
author Sobczyk, Aleksandros
Gottardo, Gioele
Matzoros, Christos K.
De Vita, Mirko
Skogh, Filip
Zouzias, Anastasios
Zhuang, Jiawei
author_facet Sobczyk, Aleksandros
Gottardo, Gioele
Matzoros, Christos K.
De Vita, Mirko
Skogh, Filip
Zouzias, Anastasios
Zhuang, Jiawei
contents Linear attention has emerged as a cornerstone for efficient long-context architectures, as evidenced by its integration into state-of-the-art open-source models including Qwen3.5/3.6, Kimi Linear, and RWKV-7. Models that incorporate linear attention layers with the so-called Delta-Rule involve the inversion of triangular matrices as a core sub-routine. This operation often forms a performance bottleneck, and, due to its high-sensitivity to numerical errors, it can significantly deteriorate end-to-end model accuracy if it is not carefully implemented. This work provides a systematic analysis of both direct and iterative triangular inversion algorithms, targeting methods that are rich in matrix products, and, therefore, have the potential to efficiently utilize modern hardware. To that end, our analysis covers a broad spectrum of mathematical and practical aspects, with a heavy focus on numerical stability, computational complexity, and, ultimately, hardware efficiency and practical considerations. We provide a rigorous experimental evaluation to verify these properties in practical scenarios, and in low-precision floating-point representations, highlighting the strengths and limitations of each method. Performance benchmarks on NPUs reveal up to $4.3\times$ speed-up against the state-of-the-art implementations of SGLang for triangular matrix inversion, leading to significant performance improvements on the entire layer level, while maintaining full end-to-end model accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21325
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers
Sobczyk, Aleksandros
Gottardo, Gioele
Matzoros, Christos K.
De Vita, Mirko
Skogh, Filip
Zouzias, Anastasios
Zhuang, Jiawei
Machine Learning
Linear attention has emerged as a cornerstone for efficient long-context architectures, as evidenced by its integration into state-of-the-art open-source models including Qwen3.5/3.6, Kimi Linear, and RWKV-7. Models that incorporate linear attention layers with the so-called Delta-Rule involve the inversion of triangular matrices as a core sub-routine. This operation often forms a performance bottleneck, and, due to its high-sensitivity to numerical errors, it can significantly deteriorate end-to-end model accuracy if it is not carefully implemented. This work provides a systematic analysis of both direct and iterative triangular inversion algorithms, targeting methods that are rich in matrix products, and, therefore, have the potential to efficiently utilize modern hardware. To that end, our analysis covers a broad spectrum of mathematical and practical aspects, with a heavy focus on numerical stability, computational complexity, and, ultimately, hardware efficiency and practical considerations. We provide a rigorous experimental evaluation to verify these properties in practical scenarios, and in low-precision floating-point representations, highlighting the strengths and limitations of each method. Performance benchmarks on NPUs reveal up to $4.3\times$ speed-up against the state-of-the-art implementations of SGLang for triangular matrix inversion, leading to significant performance improvements on the entire layer level, while maintaining full end-to-end model accuracy.
title Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers
topic Machine Learning
url https://arxiv.org/abs/2605.21325