Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wen, Yunshi, Gifford, Wesley M., Reddy, Chandra, Nguyen, Lam M., Kalagnanam, Jayant, Julius, Anak Agung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914311269515264
author Wen, Yunshi
Gifford, Wesley M.
Reddy, Chandra
Nguyen, Lam M.
Kalagnanam, Jayant
Julius, Anak Agung
author_facet Wen, Yunshi
Gifford, Wesley M.
Reddy, Chandra
Nguyen, Lam M.
Kalagnanam, Jayant
Julius, Anak Agung
contents The recent surge in Time Series Foundation Models has rapidly advanced the field, yet the heterogeneous training setups across studies make it difficult to attribute improvements to architectural innovations versus data engineering. In this work, we investigate the potential of a standard patch Transformer, demonstrating that this generic architecture achieves state-of-the-art zero-shot forecasting performance using a straightforward training protocol. We conduct a comprehensive ablation study that covers model scaling, data composition, and training techniques to isolate the essential ingredients for high performance. Our findings identify the key drivers of performance, while confirming that the generic architecture itself demonstrates excellent scalability. By strictly controlling these variables, we provide comprehensive empirical results on model scaling across multiple dimensions. We release our open-source model and detailed findings to establish a transparent, reproducible baseline for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06909
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models
Wen, Yunshi
Gifford, Wesley M.
Reddy, Chandra
Nguyen, Lam M.
Kalagnanam, Jayant
Julius, Anak Agung
Machine Learning
The recent surge in Time Series Foundation Models has rapidly advanced the field, yet the heterogeneous training setups across studies make it difficult to attribute improvements to architectural innovations versus data engineering. In this work, we investigate the potential of a standard patch Transformer, demonstrating that this generic architecture achieves state-of-the-art zero-shot forecasting performance using a straightforward training protocol. We conduct a comprehensive ablation study that covers model scaling, data composition, and training techniques to isolate the essential ingredients for high performance. Our findings identify the key drivers of performance, while confirming that the generic architecture itself demonstrates excellent scalability. By strictly controlling these variables, we provide comprehensive empirical results on model scaling across multiple dimensions. We release our open-source model and detailed findings to establish a transparent, reproducible baseline for future research.
title Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models
topic Machine Learning
url https://arxiv.org/abs/2602.06909