Fourier-Mixed Window Attention: Accelerating Informer for Long Sequence Time-Series Forecasting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tran, Nhat Thanh, Xin, Jack
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909172164984832
author Tran, Nhat Thanh
Xin, Jack
author_facet Tran, Nhat Thanh
Xin, Jack
contents We study a fast local-global window-based attention method to accelerate Informer for long sequence time-series forecasting. While window attention being local is a considerable computational saving, it lacks the ability to capture global token information which is compensated by a subsequent Fourier transform block. Our method, named FWin, does not rely on query sparsity hypothesis and an empirical approximation underlying the ProbSparse attention of Informer. Through experiments on univariate and multivariate datasets, we show that FWin transformers improve the overall prediction accuracies of Informer while accelerating its inference speeds by 1.6 to 2 times. We also provide a mathematical definition of FWin attention, and prove that it is equivalent to the canonical full attention under the block diagonal invertibility (BDI) condition of the attention matrix. The BDI is shown experimentally to hold with high probability for typical benchmark datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2307_00493
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Fourier-Mixed Window Attention: Accelerating Informer for Long Sequence Time-Series Forecasting
Tran, Nhat Thanh
Xin, Jack
Machine Learning
Artificial Intelligence
We study a fast local-global window-based attention method to accelerate Informer for long sequence time-series forecasting. While window attention being local is a considerable computational saving, it lacks the ability to capture global token information which is compensated by a subsequent Fourier transform block. Our method, named FWin, does not rely on query sparsity hypothesis and an empirical approximation underlying the ProbSparse attention of Informer. Through experiments on univariate and multivariate datasets, we show that FWin transformers improve the overall prediction accuracies of Informer while accelerating its inference speeds by 1.6 to 2 times. We also provide a mathematical definition of FWin attention, and prove that it is equivalent to the canonical full attention under the block diagonal invertibility (BDI) condition of the attention matrix. The BDI is shown experimentally to hold with high probability for typical benchmark datasets.
title Fourier-Mixed Window Attention: Accelerating Informer for Long Sequence Time-Series Forecasting
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2307.00493