Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long Sequences
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zicheng, Li, Siyuan, Wang, Li, Wang, Zedong, Liu, Yunfan, Li, Stan Z. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongVQ: Long Sequence Modeling with Vector Quantization on Structured Memory
by: Liu, Zicheng, et al.
Published: (2024)
by: Liu, Zicheng, et al.
Published: (2024)
SemiReward: A General Reward Model for Semi-supervised Learning
by: Li, Siyuan, et al.
Published: (2023)
by: Li, Siyuan, et al.
Published: (2023)
Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions
by: Tan, Cheng, et al.
Published: (2024)
by: Tan, Cheng, et al.
Published: (2024)
Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
A Survey on Mixup Augmentations and Beyond
by: Jin, Xin, et al.
Published: (2024)
by: Jin, Xin, et al.
Published: (2024)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
by: Zhou, Jingbo, et al.
Published: (2026)
by: Zhou, Jingbo, et al.
Published: (2026)
Life-Code: Central Dogma Modeling with Multi-Omics Sequence Unification
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
Switch EMA: A Free Lunch for Better Flatness and Sharpness
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
Taming LLMs by Scaling Learning Rates with Gradient Grouping
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
GenURL: A General Framework for Unsupervised Representation Learning
by: Li, Siyuan, et al.
Published: (2021)
by: Li, Siyuan, et al.
Published: (2021)
Learning Long Sequences in Spiking Neural Networks
by: Stan, Matei Ioan, et al.
Published: (2023)
by: Stan, Matei Ioan, et al.
Published: (2023)
GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models
by: Liu, Zicheng, et al.
Published: (2024)
by: Liu, Zicheng, et al.
Published: (2024)
PSC-CPI: Multi-Scale Protein Sequence-Structure Contrasting for Efficient and Generalizable Compound-Protein Interaction Prediction
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
RDesign: Hierarchical Data-efficient Representation Learning for Tertiary Structure-based RNA Design
by: Tan, Cheng, et al.
Published: (2023)
by: Tan, Cheng, et al.
Published: (2023)
Discovering the Representation Bottleneck of Graph Neural Networks
by: Wu, Fang, et al.
Published: (2022)
by: Wu, Fang, et al.
Published: (2022)
Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning
by: Wang, Zedong, et al.
Published: (2025)
by: Wang, Zedong, et al.
Published: (2025)
Teach Harder, Learn Poorer: Rethinking Hard Sample Distillation for GNN-to-MLP Knowledge Distillation
by: Wu, Lirong, et al.
Published: (2024)
by: Wu, Lirong, et al.
Published: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Efficient Sparse Selective-Update RNNs for Long-Range Sequence Modeling
by: Yin, Bojian, et al.
Published: (2026)
by: Yin, Bojian, et al.
Published: (2026)
SimVPv2: Towards Simple yet Powerful Spatiotemporal Predictive Learning
by: Tan, Cheng, et al.
Published: (2022)
by: Tan, Cheng, et al.
Published: (2022)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling
by: Zhang, Dehao, et al.
Published: (2025)
by: Zhang, Dehao, et al.
Published: (2025)
Open-World Reinforcement Learning over Long Short-Term Imagination
by: Li, Jiajian, et al.
Published: (2024)
by: Li, Jiajian, et al.
Published: (2024)
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
ZeCO: Zero Communication Overhead Sequence Parallelism for Linear Attention
by: Chou, Yuhong, et al.
Published: (2025)
by: Chou, Yuhong, et al.
Published: (2025)
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
by: Luo, Qijun, et al.
Published: (2025)
by: Luo, Qijun, et al.
Published: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
CAB: Comprehensive Attention Benchmarking on Long Sequence Modeling
by: Zhang, Jun, et al.
Published: (2022)
by: Zhang, Jun, et al.
Published: (2022)
SMR: State Memory Replay for Long Sequence Modeling
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
MergeDNA: Context-aware Genome Modeling with Dynamic Tokenization through Token Merging
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
by: Zhao, Weilin, et al.
Published: (2025)
by: Zhao, Weilin, et al.
Published: (2025)
S$^3$Attention: Improving Long Sequence Attention with Smoothed Skeleton Sketching
by: Wang, Xue, et al.
Published: (2024)
by: Wang, Xue, et al.
Published: (2024)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
by: Zhang, Geng, et al.
Published: (2025)
by: Zhang, Geng, et al.
Published: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
An Empirical Study: Extensive Deep Temporal Point Process
by: Lin, Haitao, et al.
Published: (2021)
by: Lin, Haitao, et al.
Published: (2021)
Decoupling Long- and Short-Term Patterns in Spatiotemporal Inference
by: Hu, Junfeng, et al.
Published: (2021)
by: Hu, Junfeng, et al.
Published: (2021)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
Similar Items
-
LongVQ: Long Sequence Modeling with Vector Quantization on Structured Memory
by: Liu, Zicheng, et al.
Published: (2024) -
SemiReward: A General Reward Model for Semi-supervised Learning
by: Li, Siyuan, et al.
Published: (2023) -
Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions
by: Tan, Cheng, et al.
Published: (2024) -
Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
by: Li, Siyuan, et al.
Published: (2024) -
A Survey on Mixup Augmentations and Beyond
by: Jin, Xin, et al.
Published: (2024)