Sliding Window Recurrences for Sequence Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918249684271104 |
|---|---|
| author | Secrieru, Dragos Brixi, Garyk Bengio, Yoshua Suzuki, Taiji Poli, Michael Massaroli, Stefano |
| author_facet | Secrieru, Dragos Brixi, Garyk Bengio, Yoshua Suzuki, Taiji Poli, Michael Massaroli, Stefano |
| contents | Multi-hybrid architectures are poised to take over language modeling due to better quality and performance. We introduce a hierarchical decomposition framework for linear recurrences that allows us to develop algorithms aligned with GPU memory hierarchies, yielding Sliding Window Recurrences. We focus specifically on truncating recurrences to hardware-aligned windows which are naturally jagged, limiting costly inter-warp communication. Using SWR, we develop Phalanx layers that serve as drop-in replacements for windowed attention or linear recurrences. In 1B parameter multi-hybrid models, Phalanx achieves over 10-40% speedup across 4K to 32K context length over optimized Transformers while matching perplexity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_13921 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Sliding Window Recurrences for Sequence Models Secrieru, Dragos Brixi, Garyk Bengio, Yoshua Suzuki, Taiji Poli, Michael Massaroli, Stefano Machine Learning Multi-hybrid architectures are poised to take over language modeling due to better quality and performance. We introduce a hierarchical decomposition framework for linear recurrences that allows us to develop algorithms aligned with GPU memory hierarchies, yielding Sliding Window Recurrences. We focus specifically on truncating recurrences to hardware-aligned windows which are naturally jagged, limiting costly inter-warp communication. Using SWR, we develop Phalanx layers that serve as drop-in replacements for windowed attention or linear recurrences. In 1B parameter multi-hybrid models, Phalanx achieves over 10-40% speedup across 4K to 32K context length over optimized Transformers while matching perplexity. |
| title | Sliding Window Recurrences for Sequence Models |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2512.13921 |