ParallelFlow: Parallelizing Linear Transformers via Flow Discretization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cirone, Nicola Muca, Salvi, Cristopher
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909560387665920
author Cirone, Nicola Muca
Salvi, Cristopher
author_facet Cirone, Nicola Muca
Salvi, Cristopher
contents We present a theoretical framework for analyzing linear attention models through matrix-valued state space models (SSMs). Our approach, Parallel Flows, provides a perspective that systematically decouples temporal dynamics from implementation constraints, enabling independent analysis of critical algorithmic components: chunking, parallelization, and information aggregation. Central to this framework is the reinterpretation of chunking procedures as computations of the flows governing system dynamics. This connection establishes a bridge to mathematical tools from rough path theory, opening the door to new insights into sequence modeling architectures. As a concrete application, we analyze DeltaNet in a generalized low-rank setting motivated by recent theoretical advances. Our methods allow us to design simple, streamlined generalizations of hardware-efficient algorithms present in the literature, and to provide completely different ones, inspired by rough paths techniques, with provably lower complexity. This dual contribution demonstrates how principled theoretical analysis can both explain existing practical methods and inspire fundamentally new computational approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ParallelFlow: Parallelizing Linear Transformers via Flow Discretization
Cirone, Nicola Muca
Salvi, Cristopher
Machine Learning
Dynamical Systems
We present a theoretical framework for analyzing linear attention models through matrix-valued state space models (SSMs). Our approach, Parallel Flows, provides a perspective that systematically decouples temporal dynamics from implementation constraints, enabling independent analysis of critical algorithmic components: chunking, parallelization, and information aggregation. Central to this framework is the reinterpretation of chunking procedures as computations of the flows governing system dynamics. This connection establishes a bridge to mathematical tools from rough path theory, opening the door to new insights into sequence modeling architectures. As a concrete application, we analyze DeltaNet in a generalized low-rank setting motivated by recent theoretical advances. Our methods allow us to design simple, streamlined generalizations of hardware-efficient algorithms present in the literature, and to provide completely different ones, inspired by rough paths techniques, with provably lower complexity. This dual contribution demonstrates how principled theoretical analysis can both explain existing practical methods and inspire fundamentally new computational approaches.
title ParallelFlow: Parallelizing Linear Transformers via Flow Discretization
topic Machine Learning
Dynamical Systems
url https://arxiv.org/abs/2504.00492