Not All Data are Good Labels: On the Self-supervised Labeling for Time Series Forecasting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Yuxuan, Zhang, Dalin, Liang, Yuxuan, Lu, Hua, Chen, Gang, Li, Huan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917147891990528
author Yang, Yuxuan
Zhang, Dalin
Liang, Yuxuan
Lu, Hua
Chen, Gang
Li, Huan
author_facet Yang, Yuxuan
Zhang, Dalin
Liang, Yuxuan
Lu, Hua
Chen, Gang
Li, Huan
contents Time Series Forecasting (TSF) is a crucial task in various domains, yet existing TSF models rely heavily on high-quality data and insufficiently exploit all available data. This paper explores a novel self-supervised approach to re-label time series datasets by inherently constructing candidate datasets. During the optimization of a simple reconstruction network, intermediates are used as pseudo labels in a self-supervised paradigm, improving generalization for any predictor. We introduce the Self-Correction with Adaptive Mask (SCAM), which discards overfitted components and selectively replaces them with pseudo labels generated from reconstructions. Additionally, we incorporate Spectral Norm Regularization (SNR) to further suppress overfitting from a loss landscape perspective. Our experiments on eleven real-world datasets demonstrate that SCAM consistently improves the performance of various backbone models. This work offers a new perspective on constructing datasets and enhancing the generalization of TSF models through self-supervised learning. The code is available at https://github.com/SuDIS-ZJU/SCAM.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14704
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Not All Data are Good Labels: On the Self-supervised Labeling for Time Series Forecasting
Yang, Yuxuan
Zhang, Dalin
Liang, Yuxuan
Lu, Hua
Chen, Gang
Li, Huan
Machine Learning
Artificial Intelligence
Time Series Forecasting (TSF) is a crucial task in various domains, yet existing TSF models rely heavily on high-quality data and insufficiently exploit all available data. This paper explores a novel self-supervised approach to re-label time series datasets by inherently constructing candidate datasets. During the optimization of a simple reconstruction network, intermediates are used as pseudo labels in a self-supervised paradigm, improving generalization for any predictor. We introduce the Self-Correction with Adaptive Mask (SCAM), which discards overfitted components and selectively replaces them with pseudo labels generated from reconstructions. Additionally, we incorporate Spectral Norm Regularization (SNR) to further suppress overfitting from a loss landscape perspective. Our experiments on eleven real-world datasets demonstrate that SCAM consistently improves the performance of various backbone models. This work offers a new perspective on constructing datasets and enhancing the generalization of TSF models through self-supervised learning. The code is available at https://github.com/SuDIS-ZJU/SCAM.
title Not All Data are Good Labels: On the Self-supervised Labeling for Time Series Forecasting
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.14704