MICA: Multivariate Infini Compressive Attention for Time Series Forecasting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Potosnak, Willa, Żukowska, Nina, Wiliński, Michał, Howarth, Dan, Stępka, Ignacy, Goswami, Mononito, Dubrawski, Artur
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911668947124224
author Potosnak, Willa
Żukowska, Nina
Wiliński, Michał
Howarth, Dan
Stępka, Ignacy
Goswami, Mononito
Dubrawski, Artur
author_facet Potosnak, Willa
Żukowska, Nina
Wiliński, Michał
Howarth, Dan
Stępka, Ignacy
Goswami, Mononito
Dubrawski, Artur
contents Multivariate forecasting with Transformers faces a core scalability challenge: modeling cross-channel dependencies via attention compounds attention's quadratic sequence complexity with quadratic channel scaling, making full cross-channel attention impractical for high-dimensional time series. We propose Multivariate Infini Compressive Attention (MICA), an architectural design to extend channel-independent Transformers to channel-dependent forecasting. By adapting efficient attention techniques from the sequence dimension to the channel dimension, MICA adds a cross-channel attention mechanism to channel-independent backbones that scales linearly with channel count and context length. We evaluate channel-independent Transformer architectures with and without MICA across multiple forecasting benchmarks. MICA reduces forecast error over its channel-independent counterparts by 5.4% on average and up to 25.4% on individual datasets, highlighting the importance of explicit cross-channel modeling. Moreover, models with MICA rank first among deep multivariate Transformer and MLP baselines. MICA models also scale more efficiently with respect to both channel count and context length than Transformer baselines that compute attention across both the temporal and channel dimensions, establishing compressive attention as a practical solution for scalable multivariate forecasting.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06473
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MICA: Multivariate Infini Compressive Attention for Time Series Forecasting
Potosnak, Willa
Żukowska, Nina
Wiliński, Michał
Howarth, Dan
Stępka, Ignacy
Goswami, Mononito
Dubrawski, Artur
Machine Learning
Multivariate forecasting with Transformers faces a core scalability challenge: modeling cross-channel dependencies via attention compounds attention's quadratic sequence complexity with quadratic channel scaling, making full cross-channel attention impractical for high-dimensional time series. We propose Multivariate Infini Compressive Attention (MICA), an architectural design to extend channel-independent Transformers to channel-dependent forecasting. By adapting efficient attention techniques from the sequence dimension to the channel dimension, MICA adds a cross-channel attention mechanism to channel-independent backbones that scales linearly with channel count and context length. We evaluate channel-independent Transformer architectures with and without MICA across multiple forecasting benchmarks. MICA reduces forecast error over its channel-independent counterparts by 5.4% on average and up to 25.4% on individual datasets, highlighting the importance of explicit cross-channel modeling. Moreover, models with MICA rank first among deep multivariate Transformer and MLP baselines. MICA models also scale more efficiently with respect to both channel count and context length than Transformer baselines that compute attention across both the temporal and channel dimensions, establishing compressive attention as a practical solution for scalable multivariate forecasting.
title MICA: Multivariate Infini Compressive Attention for Time Series Forecasting
topic Machine Learning
url https://arxiv.org/abs/2604.06473