Parallelizing Linear Transformers with the Delta Rule over Sequence Length
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Songlin, Wang, Bailin, Zhang, Yu, Shen, Yikang, Kim, Yoon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Gated Delta Networks: Improving Mamba2 with Delta Rule
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
In-Context Language Learning: Architectures and Algorithms
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
Learning to Decode Collaboratively with Multiple Language Models
by: Shen, Shannon Zejiang, et al.
Published: (2024)
by: Shen, Shannon Zejiang, et al.
Published: (2024)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
by: Zhou, Chenyu, et al.
Published: (2026)
by: Zhou, Chenyu, et al.
Published: (2026)
Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths
by: Ma, Xuezhe, et al.
Published: (2026)
by: Ma, Xuezhe, et al.
Published: (2026)
Diversity Measurement and Subset Selection for Instruction Tuning Datasets
by: Wang, Peiqi, et al.
Published: (2024)
by: Wang, Peiqi, et al.
Published: (2024)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
Structured Code Representations Enable Data-Efficient Adaptation of Code Language Models
by: Agarwal, Mayank, et al.
Published: (2024)
by: Agarwal, Mayank, et al.
Published: (2024)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025)
by: Nrusimha, Aniruddha, et al.
Published: (2025)
FourierNAT: A Fourier-Mixing-Based Non-Autoregressive Transformer for Parallel Sequence Generation
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
Hyperloop Transformers
by: Zeitoun, Abbas, et al.
Published: (2026)
by: Zeitoun, Abbas, et al.
Published: (2026)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
by: Shen, Shuaijie, et al.
Published: (2024)
by: Shen, Shuaijie, et al.
Published: (2024)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
by: Badger, Benjamin L.
Published: (2026)
by: Badger, Benjamin L.
Published: (2026)
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
by: Pouransari, Hadi, et al.
Published: (2024)
by: Pouransari, Hadi, et al.
Published: (2024)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
by: Reddy, Natesh, et al.
Published: (2025)
by: Reddy, Natesh, et al.
Published: (2025)
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
by: Zou, Haosheng, et al.
Published: (2025)
by: Zou, Haosheng, et al.
Published: (2025)
On the Duality between Gradient Transformations and Adapters
by: Torroba-Hennigen, Lucas, et al.
Published: (2025)
by: Torroba-Hennigen, Lucas, et al.
Published: (2025)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
by: Mao, Hanyi, et al.
Published: (2025)
by: Mao, Hanyi, et al.
Published: (2025)
R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging
by: Keita, Mamadou K., et al.
Published: (2025)
by: Keita, Mamadou K., et al.
Published: (2025)
Learning Extrapolative Sequence Transformations from Markov Chains
by: Hager, Sophia, et al.
Published: (2025)
by: Hager, Sophia, et al.
Published: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Extracting Rule-based Descriptions of Attention Features in Transformers
by: Friedman, Dan, et al.
Published: (2025)
by: Friedman, Dan, et al.
Published: (2025)
FlowBot: Inducing LLM Workflows with Bilevel Optimization and Textual Gradients
by: Yu, Hongyeon, et al.
Published: (2026)
by: Yu, Hongyeon, et al.
Published: (2026)
From Interpolation to Extrapolation: Complete Length Generalization for Arithmetic Transformers
by: Duan, Shaoxiong, et al.
Published: (2023)
by: Duan, Shaoxiong, et al.
Published: (2023)
Gated Slot Attention for Efficient Linear-Time Sequence Modeling
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
by: Zhou, Yongchao, et al.
Published: (2024)
by: Zhou, Yongchao, et al.
Published: (2024)
Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling
by: Acharya, Rishiraj
Published: (2025)
by: Acharya, Rishiraj
Published: (2025)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
by: Marsden, Annie, et al.
Published: (2024)
by: Marsden, Annie, et al.
Published: (2024)
Fusion Matters: Length-Aware Analysis of Positional-Encoding Fusion in Transformers
by: Hallam, Mohamed Amine, et al.
Published: (2026)
by: Hallam, Mohamed Amine, et al.
Published: (2026)
MoM: Linear Sequence Modeling with Mixture-of-Memories
by: Du, Jusen, et al.
Published: (2025)
by: Du, Jusen, et al.
Published: (2025)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Learning and Transferring Sparse Contextual Bigrams with Linear Transformers
by: Ren, Yunwei, et al.
Published: (2024)
by: Ren, Yunwei, et al.
Published: (2024)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
by: Sun, Zhiqing, et al.
Published: (2024)
by: Sun, Zhiqing, et al.
Published: (2024)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
by: Cheng, Zicong, et al.
Published: (2026)
by: Cheng, Zicong, et al.
Published: (2026)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
by: Ma, Xuezhe, et al.
Published: (2024)
by: Ma, Xuezhe, et al.
Published: (2024)
Similar Items
-
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023) -
Gated Delta Networks: Improving Mamba2 with Delta Rule
by: Yang, Songlin, et al.
Published: (2024) -
PaTH Attention: Position Encoding via Accumulating Householder Transformations
by: Yang, Songlin, et al.
Published: (2025) -
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024) -
In-Context Language Learning: Architectures and Algorithms
by: Akyürek, Ekin, et al.
Published: (2024)