Linear Attention Sequence Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Weigao, Qin, Zhen, Li, Dong, Shen, Xuyang, Qiao, Yu, Zhong, Yiran |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Unlocking the Secrets of Linear Complexity Sequence Model from A Unified Perspective
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Scaling Laws for Linear Complexity Language Models
by: Shen, Xuyang, et al.
Published: (2024)
by: Shen, Xuyang, et al.
Published: (2024)
Elucidating the Design Space of Decay in Linear Attention
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
HGRN2: Gated Linear RNNs with State Expansion
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Native Hybrid Attention for Efficient Sequence Modeling
by: Du, Jusen, et al.
Published: (2025)
by: Du, Jusen, et al.
Published: (2025)
CO2: Efficient Distributed Training with Full Communication-Computation Overlap
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
MoM: Linear Sequence Modeling with Mixture-of-Memories
by: Du, Jusen, et al.
Published: (2025)
by: Du, Jusen, et al.
Published: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
by: Sun, Weigao, et al.
Published: (2025)
by: Sun, Weigao, et al.
Published: (2025)
FlashSampling: Fast and Memory-Efficient Exact Sampling
by: Ruiz, Tomas, et al.
Published: (2026)
by: Ruiz, Tomas, et al.
Published: (2026)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
Liger: Linearizing Large Language Models to Gated Recurrent Structures
by: Lan, Disen, et al.
Published: (2025)
by: Lan, Disen, et al.
Published: (2025)
TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer
by: Qin, Zhen, et al.
Published: (2023)
by: Qin, Zhen, et al.
Published: (2023)
Higher-order Linear Attention
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
by: Ma, Yan, et al.
Published: (2025)
by: Ma, Yan, et al.
Published: (2025)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
by: Dong, Yihe, et al.
Published: (2025)
by: Dong, Yihe, et al.
Published: (2025)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
by: Xiao, Kaicheng, et al.
Published: (2026)
by: Xiao, Kaicheng, et al.
Published: (2026)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
by: Shen, Shuaijie, et al.
Published: (2024)
by: Shen, Shuaijie, et al.
Published: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Kimi Linear: An Expressive, Efficient Attention Architecture
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
by: Badger, Benjamin L.
Published: (2026)
by: Badger, Benjamin L.
Published: (2026)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
by: Dentamaro, Vincenzo
Published: (2025)
by: Dentamaro, Vincenzo
Published: (2025)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
by: Le, Dong, et al.
Published: (2026)
by: Le, Dong, et al.
Published: (2026)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
by: Zou, Haosheng, et al.
Published: (2025)
by: Zou, Haosheng, et al.
Published: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
by: Ge, Xuyang, et al.
Published: (2024)
by: Ge, Xuyang, et al.
Published: (2024)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
by: Rodionov, Gleb, et al.
Published: (2025)
by: Rodionov, Gleb, et al.
Published: (2025)
Long-range Modeling and Processing of Multimodal Event Sequences
by: Li, Jichu, et al.
Published: (2026)
by: Li, Jichu, et al.
Published: (2026)
Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling
by: Acharya, Rishiraj
Published: (2025)
by: Acharya, Rishiraj
Published: (2025)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
by: Jin, Zehao, et al.
Published: (2026)
by: Jin, Zehao, et al.
Published: (2026)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning
by: Zhao, Shiwan, et al.
Published: (2025)
by: Zhao, Shiwan, et al.
Published: (2025)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
by: He, Zhengfu, et al.
Published: (2025)
by: He, Zhengfu, et al.
Published: (2025)
Comba: Improving Bilinear RNNs with Closed-loop Control
by: Hu, Jiaxi, et al.
Published: (2025)
by: Hu, Jiaxi, et al.
Published: (2025)
Similar Items
-
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
by: Sun, Weigao, et al.
Published: (2025) -
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
by: Qin, Zhen, et al.
Published: (2024) -
Unlocking the Secrets of Linear Complexity Sequence Model from A Unified Perspective
by: Qin, Zhen, et al.
Published: (2024) -
Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention
by: Qin, Zhen, et al.
Published: (2024) -
Scaling Laws for Linear Complexity Language Models
by: Shen, Xuyang, et al.
Published: (2024)