When Does Multimodality Lead to Better Time Series Forecasting?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xiyuan, Han, Boran, Fang, Haoyang, Ansari, Abdul Fatir, Zhang, Shuai, Maddix, Danielle C., Hu, Cuixiong, Wilson, Andrew Gordon, Mahoney, Michael W., Wang, Hao, Liu, Yan, Rangwala, Huzefa, Karypis, George, Wang, Bernie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911183973384192
author Zhang, Xiyuan
Han, Boran
Fang, Haoyang
Ansari, Abdul Fatir
Zhang, Shuai
Maddix, Danielle C.
Hu, Cuixiong
Wilson, Andrew Gordon
Mahoney, Michael W.
Wang, Hao
Liu, Yan
Rangwala, Huzefa
Karypis, George
Wang, Bernie
author_facet Zhang, Xiyuan
Han, Boran
Fang, Haoyang
Ansari, Abdul Fatir
Zhang, Shuai
Maddix, Danielle C.
Hu, Cuixiong
Wilson, Andrew Gordon
Mahoney, Michael W.
Wang, Hao
Liu, Yan
Rangwala, Huzefa
Karypis, George
Wang, Bernie
contents Recently, there has been growing interest in incorporating textual information into foundation models for time series forecasting. However, it remains unclear whether and under what conditions such multimodal integration consistently yields gains. We systematically investigate these questions across a diverse benchmark of 16 forecasting tasks spanning 7 domains, including health, environment, and economics. We evaluate two popular multimodal forecasting paradigms: aligning-based methods, which align time series and text representations; and prompting-based methods, which directly prompt large language models for forecasting. Our findings reveal that the benefits of multimodality are highly condition-dependent. While we confirm reported gains in some settings, these improvements are not universal across datasets or models. To move beyond empirical observations, we disentangle the effects of model architectural properties and data characteristics, drawing data-agnostic insights that generalize across domains. Our findings highlight that on the modeling side, incorporating text information is most helpful given (1) high-capacity text models, (2) comparatively weaker time series models, and (3) appropriate aligning strategies. On the data side, performance gains are more likely when (4) sufficient training data is available and (5) the text offers complementary predictive signal beyond what is already captured from the time series alone. Our study offers a rigorous, quantitative foundation for understanding when multimodality can be expected to aid forecasting tasks, and reveals that its benefits are neither universal nor always aligned with intuition.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21611
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Does Multimodality Lead to Better Time Series Forecasting?
Zhang, Xiyuan
Han, Boran
Fang, Haoyang
Ansari, Abdul Fatir
Zhang, Shuai
Maddix, Danielle C.
Hu, Cuixiong
Wilson, Andrew Gordon
Mahoney, Michael W.
Wang, Hao
Liu, Yan
Rangwala, Huzefa
Karypis, George
Wang, Bernie
Computation and Language
Artificial Intelligence
Machine Learning
Recently, there has been growing interest in incorporating textual information into foundation models for time series forecasting. However, it remains unclear whether and under what conditions such multimodal integration consistently yields gains. We systematically investigate these questions across a diverse benchmark of 16 forecasting tasks spanning 7 domains, including health, environment, and economics. We evaluate two popular multimodal forecasting paradigms: aligning-based methods, which align time series and text representations; and prompting-based methods, which directly prompt large language models for forecasting. Our findings reveal that the benefits of multimodality are highly condition-dependent. While we confirm reported gains in some settings, these improvements are not universal across datasets or models. To move beyond empirical observations, we disentangle the effects of model architectural properties and data characteristics, drawing data-agnostic insights that generalize across domains. Our findings highlight that on the modeling side, incorporating text information is most helpful given (1) high-capacity text models, (2) comparatively weaker time series models, and (3) appropriate aligning strategies. On the data side, performance gains are more likely when (4) sufficient training data is available and (5) the text offers complementary predictive signal beyond what is already captured from the time series alone. Our study offers a rigorous, quantitative foundation for understanding when multimodality can be expected to aid forecasting tasks, and reveals that its benefits are neither universal nor always aligned with intuition.
title When Does Multimodality Lead to Better Time Series Forecasting?
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.21611