Enhancing Linear Attention with Residual Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lai, Xunhao, Kang, Jialiang, Lu, Jianqiao, Lin, Tong, Zhao, Pengyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
Attention Sinks and Outliers in Attention Residuals
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
KVBuffer: IO-aware Serving for Linear Attention
von: Zou, Longwei, et al.
Veröffentlicht: (2026)
von: Zou, Longwei, et al.
Veröffentlicht: (2026)
Theory Foundation of Physics-Enhanced Residual Learning
von: Liang, Shixiao, et al.
Veröffentlicht: (2025)
von: Liang, Shixiao, et al.
Veröffentlicht: (2025)
Drug Synergy Prediction via Residual Graph Isomorphism Networks and Attention Mechanisms
von: Song, Jiyan, et al.
Veröffentlicht: (2026)
von: Song, Jiyan, et al.
Veröffentlicht: (2026)
Exact Linear Attention
von: Ou, Weinuo
Veröffentlicht: (2026)
von: Ou, Weinuo
Veröffentlicht: (2026)
Kaczmarz Linear Attention
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting
von: Lu, Jiecheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2025)
ResidualDroppath: Enhancing Feature Reuse over Residual Connections
von: Park, Sejik
Veröffentlicht: (2024)
von: Park, Sejik
Veröffentlicht: (2024)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
von: Zuo, Yifei, et al.
Veröffentlicht: (2025)
von: Zuo, Yifei, et al.
Veröffentlicht: (2025)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
von: Hu, Wenjie, et al.
Veröffentlicht: (2025)
von: Hu, Wenjie, et al.
Veröffentlicht: (2025)
E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory
von: Huang, Lin, et al.
Veröffentlicht: (2026)
von: Huang, Lin, et al.
Veröffentlicht: (2026)
Unsupervised Learning Method for the Wave Equation Based on Finite Difference Residual Constraints Loss
von: Feng, Xin, et al.
Veröffentlicht: (2024)
von: Feng, Xin, et al.
Veröffentlicht: (2024)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
von: Wi, Hyowon, et al.
Veröffentlicht: (2025)
von: Wi, Hyowon, et al.
Veröffentlicht: (2025)
LNUCB-TA: Linear-nonlinear Hybrid Bandit Learning with Temporal Attention
von: Khosravi, Hamed, et al.
Veröffentlicht: (2025)
von: Khosravi, Hamed, et al.
Veröffentlicht: (2025)
AICRN: Attention-Integrated Convolutional Residual Network for Interpretable Electrocardiogram Analysis
von: Jayakody, J. M. I. H., et al.
Veröffentlicht: (2025)
von: Jayakody, J. M. I. H., et al.
Veröffentlicht: (2025)
GPT Carry-On: Training Foundation Model for Customization Could Be Simple, Scalable and Affordable
von: Wangni, Jianqiao
Veröffentlicht: (2025)
von: Wangni, Jianqiao
Veröffentlicht: (2025)
Interpreting Machine Learning Models for Room Temperature Prediction in Non-domestic Buildings
von: Mao, Jianqiao, et al.
Veröffentlicht: (2021)
von: Mao, Jianqiao, et al.
Veröffentlicht: (2021)
RDIT: Residual-based Diffusion Implicit Models for Probabilistic Time Series Forecasting
von: Lai, Chih-Yu, et al.
Veröffentlicht: (2025)
von: Lai, Chih-Yu, et al.
Veröffentlicht: (2025)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022)
DGTN: Graph-Enhanced Transformer with Diffusive Attention Gating Mechanism for Enzyme DDG Prediction
von: Lin, Abigail
Veröffentlicht: (2025)
von: Lin, Abigail
Veröffentlicht: (2025)
Linear Attention for Efficient Bidirectional Sequence Modeling
von: Afzal, Arshia, et al.
Veröffentlicht: (2025)
von: Afzal, Arshia, et al.
Veröffentlicht: (2025)
In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
von: Goel, Ayush, et al.
Veröffentlicht: (2026)
von: Goel, Ayush, et al.
Veröffentlicht: (2026)
Adaptive Memory Decay for Log-Linear Attention
von: Amin, Yaxita, et al.
Veröffentlicht: (2026)
von: Amin, Yaxita, et al.
Veröffentlicht: (2026)
Linear Attention is Enough in Spatial-Temporal Forecasting
von: Ning, Xinyu
Veröffentlicht: (2024)
von: Ning, Xinyu
Veröffentlicht: (2024)
State Rank Dynamics in Linear Attention LLMs
von: Sun, Ao, et al.
Veröffentlicht: (2026)
von: Sun, Ao, et al.
Veröffentlicht: (2026)
A New Perspective on Time Series Anomaly Detection: Faster Patch-based Broad Learning System
von: Li, Pengyu, et al.
Veröffentlicht: (2024)
von: Li, Pengyu, et al.
Veröffentlicht: (2024)
DRSLF: Double Regularized Second-Order Low-Rank Representation for Web Service QoS Prediction
von: Wu, Hao, et al.
Veröffentlicht: (2025)
von: Wu, Hao, et al.
Veröffentlicht: (2025)
Attention-Aided MMSE for OFDM Channel Estimation: Learning Linear Filters with Attention
von: Ha, TaeJun, et al.
Veröffentlicht: (2025)
von: Ha, TaeJun, et al.
Veröffentlicht: (2025)
Mask-PINNs: Mitigating Internal Covariate Shift in Physics-Informed Neural Networks
von: Jiang, Feilong, et al.
Veröffentlicht: (2025)
von: Jiang, Feilong, et al.
Veröffentlicht: (2025)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
von: Beck, Maximilian, et al.
Veröffentlicht: (2025)
von: Beck, Maximilian, et al.
Veröffentlicht: (2025)
Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
von: Dentamaro, Vincenzo
Veröffentlicht: (2025)
von: Dentamaro, Vincenzo
Veröffentlicht: (2025)
Attention-Enhanced Deep Learning for Device-Free Through-the-Wall Presence Detection Using Indoor WiFi Systems
von: Shen, Li-Hsiang, et al.
Veröffentlicht: (2023)
von: Shen, Li-Hsiang, et al.
Veröffentlicht: (2023)
Beyond Linearity in Attention Projections: The Case for Nonlinear Queries
von: Karbevski, Marko
Veröffentlicht: (2026)
von: Karbevski, Marko
Veröffentlicht: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
Residual GRU+MHSA: A Lightweight Hybrid Recurrent Attention Model for Cardiovascular Disease Detection
von: Dash, Tejaswani, et al.
Veröffentlicht: (2025)
von: Dash, Tejaswani, et al.
Veröffentlicht: (2025)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
von: Joshi, Sahil, et al.
Veröffentlicht: (2025)
von: Joshi, Sahil, et al.
Veröffentlicht: (2025)
CYCLE: Cross-Year Contrastive Learning in Entity-Linking
von: Zhang, Pengyu, et al.
Veröffentlicht: (2024)
von: Zhang, Pengyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
von: Lai, Xunhao, et al.
Veröffentlicht: (2025) -
Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?
von: Chen, Xi, et al.
Veröffentlicht: (2024) -
Attention Sinks and Outliers in Attention Residuals
von: Luo, Haozheng, et al.
Veröffentlicht: (2026) -
KVBuffer: IO-aware Serving for Linear Attention
von: Zou, Longwei, et al.
Veröffentlicht: (2026) -
Theory Foundation of Physics-Enhanced Residual Learning
von: Liang, Shixiao, et al.
Veröffentlicht: (2025)