Investigating Compositional Reasoning in Time Series Foundation Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Potosnak, Willa, Challu, Cristian, Goswami, Mononito, Olivares, Kin G., Wiliński, Michał, Żukowska, Nina, Dubrawski, Artur
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911146508812288
author Potosnak, Willa
Challu, Cristian
Goswami, Mononito
Olivares, Kin G.
Wiliński, Michał
Żukowska, Nina
Dubrawski, Artur
author_facet Potosnak, Willa
Challu, Cristian
Goswami, Mononito
Olivares, Kin G.
Wiliński, Michał
Żukowska, Nina
Dubrawski, Artur
contents Large pre-trained time series foundation models (TSFMs) have demonstrated promising zero-shot performance across a wide range of domains. However, a question remains: Do TSFMs succeed by memorizing patterns in training data, or do they possess the ability to reason about such patterns? While reasoning is a topic of great interest in the study of Large Language Models (LLMs), it is undefined and largely unexplored in the context of TSFMs. In this work, inspired by language modeling literature, we formally define compositional reasoning in forecasting and distinguish it from in-distribution generalization. We evaluate the reasoning and generalization capabilities of 16 popular deep learning forecasting models on multiple synthetic and real-world datasets. Additionally, through controlled studies, we systematically examine which design choices in 7 popular open-source TSFMs contribute to improved reasoning capabilities. Our study yields key insights into the impact of TSFM architecture design on compositional reasoning and generalization. We find that patch-based Transformers have the best reasoning performance, closely followed by residualized MLP-based architectures, which are 97\% less computationally complex in terms of FLOPs and 86\% smaller in terms of the number of trainable parameters. Interestingly, in some zero-shot out-of-distribution scenarios, these models can outperform moving average and exponential smoothing statistical baselines trained on in-distribution data. Only a few design choices, such as the tokenization method, had a significant (negative) impact on Transformer model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2502_06037
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating Compositional Reasoning in Time Series Foundation Models
Potosnak, Willa
Challu, Cristian
Goswami, Mononito
Olivares, Kin G.
Wiliński, Michał
Żukowska, Nina
Dubrawski, Artur
Machine Learning
Large pre-trained time series foundation models (TSFMs) have demonstrated promising zero-shot performance across a wide range of domains. However, a question remains: Do TSFMs succeed by memorizing patterns in training data, or do they possess the ability to reason about such patterns? While reasoning is a topic of great interest in the study of Large Language Models (LLMs), it is undefined and largely unexplored in the context of TSFMs. In this work, inspired by language modeling literature, we formally define compositional reasoning in forecasting and distinguish it from in-distribution generalization. We evaluate the reasoning and generalization capabilities of 16 popular deep learning forecasting models on multiple synthetic and real-world datasets. Additionally, through controlled studies, we systematically examine which design choices in 7 popular open-source TSFMs contribute to improved reasoning capabilities. Our study yields key insights into the impact of TSFM architecture design on compositional reasoning and generalization. We find that patch-based Transformers have the best reasoning performance, closely followed by residualized MLP-based architectures, which are 97\% less computationally complex in terms of FLOPs and 86\% smaller in terms of the number of trainable parameters. Interestingly, in some zero-shot out-of-distribution scenarios, these models can outperform moving average and exponential smoothing statistical baselines trained on in-distribution data. Only a few design choices, such as the tokenization method, had a significant (negative) impact on Transformer model performance.
title Investigating Compositional Reasoning in Time Series Foundation Models
topic Machine Learning
url https://arxiv.org/abs/2502.06037