Are Self-Attentions Effective for Time Series Forecasting?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Dongbin, Park, Jinseong, Lee, Jaewook, Kim, Hoki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910759809712128
author Kim, Dongbin
Park, Jinseong
Lee, Jaewook
Kim, Hoki
author_facet Kim, Dongbin
Park, Jinseong
Lee, Jaewook
Kim, Hoki
contents Time series forecasting is crucial for applications across multiple domains and various scenarios. Although Transformer models have dramatically advanced the landscape of forecasting, their effectiveness remains debated. Recent findings have indicated that simpler linear models might outperform complex Transformer-based approaches, highlighting the potential for more streamlined architectures. In this paper, we shift the focus from evaluating the overall Transformer architecture to specifically examining the effectiveness of self-attention for time series forecasting. To this end, we introduce a new architecture, Cross-Attention-only Time Series transformer (CATS), that rethinks the traditional Transformer framework by eliminating self-attention and leveraging cross-attention mechanisms instead. By establishing future horizon-dependent parameters as queries and enhanced parameter sharing, our model not only improves long-term forecasting accuracy but also reduces the number of parameters and memory usage. Extensive experiment across various datasets demonstrates that our model achieves superior performance with the lowest mean squared error and uses fewer parameters compared to existing models. The implementation of our model is available at: https://github.com/dongbeank/CATS.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16877
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are Self-Attentions Effective for Time Series Forecasting?
Kim, Dongbin
Park, Jinseong
Lee, Jaewook
Kim, Hoki
Machine Learning
Artificial Intelligence
Time series forecasting is crucial for applications across multiple domains and various scenarios. Although Transformer models have dramatically advanced the landscape of forecasting, their effectiveness remains debated. Recent findings have indicated that simpler linear models might outperform complex Transformer-based approaches, highlighting the potential for more streamlined architectures. In this paper, we shift the focus from evaluating the overall Transformer architecture to specifically examining the effectiveness of self-attention for time series forecasting. To this end, we introduce a new architecture, Cross-Attention-only Time Series transformer (CATS), that rethinks the traditional Transformer framework by eliminating self-attention and leveraging cross-attention mechanisms instead. By establishing future horizon-dependent parameters as queries and enhanced parameter sharing, our model not only improves long-term forecasting accuracy but also reduces the number of parameters and memory usage. Extensive experiment across various datasets demonstrates that our model achieves superior performance with the lowest mean squared error and uses fewer parameters compared to existing models. The implementation of our model is available at: https://github.com/dongbeank/CATS.
title Are Self-Attentions Effective for Time Series Forecasting?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.16877