Attention as Robust Representation for Time Series Forecasting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niu, PeiSong, Zhou, Tian, Wang, Xue, Sun, Liang, Jin, Rong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909098900979712
author Niu, PeiSong
Zhou, Tian
Wang, Xue
Sun, Liang
Jin, Rong
author_facet Niu, PeiSong
Zhou, Tian
Wang, Xue
Sun, Liang
Jin, Rong
contents Time series forecasting is essential for many practical applications, with the adoption of transformer-based models on the rise due to their impressive performance in NLP and CV. Transformers' key feature, the attention mechanism, dynamically fusing embeddings to enhance data representation, often relegating attention weights to a byproduct role. Yet, time series data, characterized by noise and non-stationarity, poses significant forecasting challenges. Our approach elevates attention weights as the primary representation for time series, capitalizing on the temporal relationships among data points to improve forecasting accuracy. Our study shows that an attention map, structured using global landmarks and local windows, acts as a robust kernel representation for data points, withstanding noise and shifts in distribution. Our method outperforms state-of-the-art models, reducing mean squared error (MSE) in multivariate time series forecasting by a notable 3.6% without altering the core neural network architecture. It serves as a versatile component that can readily replace recent patching based embedding schemes in transformer-based models, boosting their performance.
format Preprint
id arxiv_https___arxiv_org_abs_2402_05370
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Attention as Robust Representation for Time Series Forecasting
Niu, PeiSong
Zhou, Tian
Wang, Xue
Sun, Liang
Jin, Rong
Machine Learning
Artificial Intelligence
Time series forecasting is essential for many practical applications, with the adoption of transformer-based models on the rise due to their impressive performance in NLP and CV. Transformers' key feature, the attention mechanism, dynamically fusing embeddings to enhance data representation, often relegating attention weights to a byproduct role. Yet, time series data, characterized by noise and non-stationarity, poses significant forecasting challenges. Our approach elevates attention weights as the primary representation for time series, capitalizing on the temporal relationships among data points to improve forecasting accuracy. Our study shows that an attention map, structured using global landmarks and local windows, acts as a robust kernel representation for data points, withstanding noise and shifts in distribution. Our method outperforms state-of-the-art models, reducing mean squared error (MSE) in multivariate time series forecasting by a notable 3.6% without altering the core neural network architecture. It serves as a versatile component that can readily replace recent patching based embedding schemes in transformer-based models, boosting their performance.
title Attention as Robust Representation for Time Series Forecasting
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2402.05370