The Power of Architecture: Deep Dive into Transformer Architectures for Long-Term Time Series Forecasting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Lefei, Chen, Mouxiang, Fu, Han, Ren, Xiaoxue, Wang, Xiaoyun Joy, Sun, Jianling, Li, Zhuo, Liu, Chenghao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918095880192000
author Shen, Lefei
Chen, Mouxiang
Fu, Han
Ren, Xiaoxue
Wang, Xiaoyun Joy
Sun, Jianling
Li, Zhuo
Liu, Chenghao
author_facet Shen, Lefei
Chen, Mouxiang
Fu, Han
Ren, Xiaoxue
Wang, Xiaoyun Joy
Sun, Jianling
Li, Zhuo
Liu, Chenghao
contents Transformer-based models have recently become dominant in Long-term Time Series Forecasting (LTSF), yet the variations in their architecture, such as encoder-only, encoder-decoder, and decoder-only designs, raise a crucial question: What Transformer architecture works best for LTSF tasks? However, existing models are often tightly coupled with various time-series-specific designs, making it difficult to isolate the impact of the architecture itself. To address this, we propose a novel taxonomy that disentangles these designs, enabling clearer and more unified comparisons of Transformer architectures. Our taxonomy considers key aspects such as attention mechanisms, forecasting aggregations, forecasting paradigms, and normalization layers. Through extensive experiments, we uncover several key insights: bi-directional attention with joint-attention is most effective; more complete forecasting aggregation improves performance; and the direct-mapping paradigm outperforms autoregressive approaches. Furthermore, our combined model, utilizing optimal architectural choices, consistently outperforms several existing models, reinforcing the validity of our conclusions. We hope these findings offer valuable guidance for future research on Transformer architectural designs in LTSF. Our code is available at https://github.com/HALF111/TSF_architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13043
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Power of Architecture: Deep Dive into Transformer Architectures for Long-Term Time Series Forecasting
Shen, Lefei
Chen, Mouxiang
Fu, Han
Ren, Xiaoxue
Wang, Xiaoyun Joy
Sun, Jianling
Li, Zhuo
Liu, Chenghao
Machine Learning
Transformer-based models have recently become dominant in Long-term Time Series Forecasting (LTSF), yet the variations in their architecture, such as encoder-only, encoder-decoder, and decoder-only designs, raise a crucial question: What Transformer architecture works best for LTSF tasks? However, existing models are often tightly coupled with various time-series-specific designs, making it difficult to isolate the impact of the architecture itself. To address this, we propose a novel taxonomy that disentangles these designs, enabling clearer and more unified comparisons of Transformer architectures. Our taxonomy considers key aspects such as attention mechanisms, forecasting aggregations, forecasting paradigms, and normalization layers. Through extensive experiments, we uncover several key insights: bi-directional attention with joint-attention is most effective; more complete forecasting aggregation improves performance; and the direct-mapping paradigm outperforms autoregressive approaches. Furthermore, our combined model, utilizing optimal architectural choices, consistently outperforms several existing models, reinforcing the validity of our conclusions. We hope these findings offer valuable guidance for future research on Transformer architectural designs in LTSF. Our code is available at https://github.com/HALF111/TSF_architecture.
title The Power of Architecture: Deep Dive into Transformer Architectures for Long-Term Time Series Forecasting
topic Machine Learning
url https://arxiv.org/abs/2507.13043