Why Attention Fails: The Degeneration of Transformers into MLPs in Time Series Forecasting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Zida, Zhu, Jiayi, Sun, Weiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Why Do Transformers Fail to Forecast Time Series In-Context?
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
Hyperparameter Tuning MLPs for Probabilistic Time Series Forecasting
von: Madhusudhanan, Kiran, et al.
Veröffentlicht: (2024)
von: Madhusudhanan, Kiran, et al.
Veröffentlicht: (2024)
Boosting MLPs with a Coarsening Strategy for Long-Term Time Series Forecasting
von: Bian, Nannan, et al.
Veröffentlicht: (2024)
von: Bian, Nannan, et al.
Veröffentlicht: (2024)
MDMLP-EIA: Multi-domain Dynamic MLPs with Energy Invariant Attention for Time Series Forecasting
von: Zhang, Hu, et al.
Veröffentlicht: (2025)
von: Zhang, Hu, et al.
Veröffentlicht: (2025)
Attention as Robust Representation for Time Series Forecasting
von: Niu, PeiSong, et al.
Veröffentlicht: (2024)
von: Niu, PeiSong, et al.
Veröffentlicht: (2024)
Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time Series
von: Yu, Guoqi, et al.
Veröffentlicht: (2026)
von: Yu, Guoqi, et al.
Veröffentlicht: (2026)
Enhancing Time Series Forecasting with Fuzzy Attention-Integrated Transformers
von: Chakraborty, Sanjay, et al.
Veröffentlicht: (2025)
von: Chakraborty, Sanjay, et al.
Veröffentlicht: (2025)
Why Model Selection Fails in Time Series Forecasting: An Empirical Study of Instability Across Data Regimes
von: Akinci, Tahir Cetin, et al.
Veröffentlicht: (2026)
von: Akinci, Tahir Cetin, et al.
Veröffentlicht: (2026)
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
von: Qiu, Haiquan, et al.
Veröffentlicht: (2025)
von: Qiu, Haiquan, et al.
Veröffentlicht: (2025)
WaveRoRA: Wavelet Rotary Route Attention for Multivariate Time Series Forecasting
von: Liang, Aobo, et al.
Veröffentlicht: (2024)
von: Liang, Aobo, et al.
Veröffentlicht: (2024)
PSformer: Parameter-efficient Transformer with Segment Attention for Time Series Forecasting
von: Wang, Yanlong, et al.
Veröffentlicht: (2024)
von: Wang, Yanlong, et al.
Veröffentlicht: (2024)
PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
von: Sun, Tian, et al.
Veröffentlicht: (2025)
von: Sun, Tian, et al.
Veröffentlicht: (2025)
LATST: Are Transformers Necessarily Complex for Time-Series Forecasting
von: Liang, Dizhen
Veröffentlicht: (2024)
von: Liang, Dizhen
Veröffentlicht: (2024)
TimeFormer: Transformer with Attention Modulation Empowered by Temporal Characteristics for Time Series Forecasting
von: Liu, Zhipeng, et al.
Veröffentlicht: (2025)
von: Liu, Zhipeng, et al.
Veröffentlicht: (2025)
Sentinel: Multi-Patch Transformer with Temporal and Channel Attention for Time Series Forecasting
von: Villaboni, Davide, et al.
Veröffentlicht: (2025)
von: Villaboni, Davide, et al.
Veröffentlicht: (2025)
Integrating Quantum-Classical Attention in Patch Transformers for Enhanced Time Series Forecasting
von: Chakraborty, Sanjay, et al.
Veröffentlicht: (2025)
von: Chakraborty, Sanjay, et al.
Veröffentlicht: (2025)
When Will It Fail?: Anomaly to Prompt for Forecasting Future Anomalies in Time Series
von: Park, Min-Yeong, et al.
Veröffentlicht: (2025)
von: Park, Min-Yeong, et al.
Veröffentlicht: (2025)
Local Attention Mechanism: Boosting the Transformer Architecture for Long-Sequence Time Series Forecasting
von: Aguilera-Martos, Ignacio, et al.
Veröffentlicht: (2024)
von: Aguilera-Martos, Ignacio, et al.
Veröffentlicht: (2024)
CAPS: Unifying Attention, Recurrence, and Alignment in Transformer-based Time Series Forecasting
von: Pati, Viresh, et al.
Veröffentlicht: (2026)
von: Pati, Viresh, et al.
Veröffentlicht: (2026)
Multi-Order Wavelet Derivative Transform for Deep Time Series Forecasting
von: Zhou, Ziyu, et al.
Veröffentlicht: (2025)
von: Zhou, Ziyu, et al.
Veröffentlicht: (2025)
S2TX: Cross-Attention Multi-Scale State-Space Transformer for Time Series Forecasting
von: Wu, Zihao, et al.
Veröffentlicht: (2025)
von: Wu, Zihao, et al.
Veröffentlicht: (2025)
Revisiting Attention for Multivariate Time Series Forecasting
von: Wu, Haixiang
Veröffentlicht: (2024)
von: Wu, Haixiang
Veröffentlicht: (2024)
Are Self-Attentions Effective for Time Series Forecasting?
von: Kim, Dongbin, et al.
Veröffentlicht: (2024)
von: Kim, Dongbin, et al.
Veröffentlicht: (2024)
DWAFM: Dynamic Weighted Graph Structure Embedding Integrated with Attention and Frequency-Domain MLPs for Traffic Forecasting
von: Shi, Sen, et al.
Veröffentlicht: (2026)
von: Shi, Sen, et al.
Veröffentlicht: (2026)
Patch-Level Tokenization with CNN Encoders and Attention for Improved Transformer Time-Series Forecasting
von: Nagrath, Saurish, et al.
Veröffentlicht: (2026)
von: Nagrath, Saurish, et al.
Veröffentlicht: (2026)
Bi-Mamba+: Bidirectional Mamba for Time Series Forecasting
von: Liang, Aobo, et al.
Veröffentlicht: (2024)
von: Liang, Aobo, et al.
Veröffentlicht: (2024)
iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
von: Liu, Yong, et al.
Veröffentlicht: (2023)
von: Liu, Yong, et al.
Veröffentlicht: (2023)
SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention
von: Ilbert, Romain, et al.
Veröffentlicht: (2024)
von: Ilbert, Romain, et al.
Veröffentlicht: (2024)
MICA: Multivariate Infini Compressive Attention for Time Series Forecasting
von: Potosnak, Willa, et al.
Veröffentlicht: (2026)
von: Potosnak, Willa, et al.
Veröffentlicht: (2026)
Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences
von: Chen, Xiwen, et al.
Veröffentlicht: (2025)
von: Chen, Xiwen, et al.
Veröffentlicht: (2025)
Addressing Concept Shift in Online Time Series Forecasting: Detect-then-Adapt
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
FAITH: Frequency-domain Attention In Two Horizons for Time Series Forecasting
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
von: Li, Ruiqi, et al.
Veröffentlicht: (2024)
Generative Pretrained Hierarchical Transformer for Time Series Forecasting
von: Liu, Zhiding, et al.
Veröffentlicht: (2024)
von: Liu, Zhiding, et al.
Veröffentlicht: (2024)
Constructing Efficient Fact-Storing MLPs for Transformers
von: Dugan, Owen, et al.
Veröffentlicht: (2025)
von: Dugan, Owen, et al.
Veröffentlicht: (2025)
DeformTime: Capturing Variable Dependencies with Deformable Attention for Time Series Forecasting
von: Shu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Shu, Yuxuan, et al.
Veröffentlicht: (2024)
MCformer: Multivariate Time Series Forecasting with Mixed-Channels Transformer
von: Han, Wenyong, et al.
Veröffentlicht: (2024)
von: Han, Wenyong, et al.
Veröffentlicht: (2024)
SST: Multi-Scale Hybrid Mamba-Transformer Experts for Time Series Forecasting
von: Xu, Xiongxiao, et al.
Veröffentlicht: (2024)
von: Xu, Xiongxiao, et al.
Veröffentlicht: (2024)
Forecasting with Guidance: Representation-Level Supervision for Time Series Forecasting
von: Wang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Wang, Jiacheng, et al.
Veröffentlicht: (2026)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
von: Zhao, Yanjun, et al.
Veröffentlicht: (2024)
von: Zhao, Yanjun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Why Do Transformers Fail to Forecast Time Series In-Context?
von: Zhou, Yufa, et al.
Veröffentlicht: (2025) -
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
von: Ke, Yekun, et al.
Veröffentlicht: (2024) -
Hyperparameter Tuning MLPs for Probabilistic Time Series Forecasting
von: Madhusudhanan, Kiran, et al.
Veröffentlicht: (2024) -
Boosting MLPs with a Coarsening Strategy for Long-Term Time Series Forecasting
von: Bian, Nannan, et al.
Veröffentlicht: (2024) -
MDMLP-EIA: Multi-domain Dynamic MLPs with Energy Invariant Attention for Time Series Forecasting
von: Zhang, Hu, et al.
Veröffentlicht: (2025)