BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shao, Zezhi, Li, Yujie, Wang, Fei, Yu, Chengqing, Fu, Yisong, Qian, Tangwen, Xu, Bin, Diao, Boyu, Xu, Yongjun, Cheng, Xueqi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909624033083392
author Shao, Zezhi
Li, Yujie
Wang, Fei
Yu, Chengqing
Fu, Yisong
Qian, Tangwen
Xu, Bin
Diao, Boyu
Xu, Yongjun
Cheng, Xueqi
author_facet Shao, Zezhi
Li, Yujie
Wang, Fei
Yu, Chengqing
Fu, Yisong
Qian, Tangwen
Xu, Bin
Diao, Boyu
Xu, Yongjun
Cheng, Xueqi
contents The advent of universal time series forecasting models has revolutionized zero-shot forecasting across diverse domains, yet the critical role of data diversity in training these models remains underexplored. Existing large-scale time series datasets often suffer from inherent biases and imbalanced distributions, leading to suboptimal model performance and generalization. To address this gap, we introduce BLAST, a novel pre-training corpus designed to enhance data diversity through a balanced sampling strategy. First, BLAST incorporates 321 billion observations from publicly available datasets and employs a comprehensive suite of statistical metrics to characterize time series patterns. Then, to facilitate pattern-oriented sampling, the data is implicitly clustered using grid-based partitioning. Furthermore, by integrating grid sampling and grid mixup techniques, BLAST ensures a balanced and representative coverage of diverse patterns. Experimental results demonstrate that models pre-trained on BLAST achieve state-of-the-art performance with a fraction of the computational resources and training tokens required by existing methods. Our findings highlight the pivotal role of data diversity in improving both training efficiency and model performance for the universal forecasting task.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17871
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models
Shao, Zezhi
Li, Yujie
Wang, Fei
Yu, Chengqing
Fu, Yisong
Qian, Tangwen
Xu, Bin
Diao, Boyu
Xu, Yongjun
Cheng, Xueqi
Machine Learning
The advent of universal time series forecasting models has revolutionized zero-shot forecasting across diverse domains, yet the critical role of data diversity in training these models remains underexplored. Existing large-scale time series datasets often suffer from inherent biases and imbalanced distributions, leading to suboptimal model performance and generalization. To address this gap, we introduce BLAST, a novel pre-training corpus designed to enhance data diversity through a balanced sampling strategy. First, BLAST incorporates 321 billion observations from publicly available datasets and employs a comprehensive suite of statistical metrics to characterize time series patterns. Then, to facilitate pattern-oriented sampling, the data is implicitly clustered using grid-based partitioning. Furthermore, by integrating grid sampling and grid mixup techniques, BLAST ensures a balanced and representative coverage of diverse patterns. Experimental results demonstrate that models pre-trained on BLAST achieve state-of-the-art performance with a fraction of the computational resources and training tokens required by existing methods. Our findings highlight the pivotal role of data diversity in improving both training efficiency and model performance for the universal forecasting task.
title BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models
topic Machine Learning
url https://arxiv.org/abs/2505.17871