Are Synthetic Time-series Data Really not as Good as Real Data?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Fanzhe, Chen, Junru, Zhang, Jing, Yang, Carl, Ma, Lvbin, Yang, Yang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911769029509120
author Fu, Fanzhe
Chen, Junru
Zhang, Jing
Yang, Carl
Ma, Lvbin
Yang, Yang
author_facet Fu, Fanzhe
Chen, Junru
Zhang, Jing
Yang, Carl
Ma, Lvbin
Yang, Yang
contents Time-series data presents limitations stemming from data quality issues, bias and vulnerabilities, and generalization problem. Integrating universal data synthesis methods holds promise in improving generalization. However, current methods cannot guarantee that the generator's output covers all unseen real data. In this paper, we introduce InfoBoost -- a highly versatile cross-domain data synthesizing framework with time series representation learning capability. We have developed a method based on synthetic data that enables model training without the need for real data, surpassing the performance of models trained with real data. Additionally, we have trained a universal feature extractor based on our synthetic data that is applicable to all time-series data. Our approach overcomes interference from multiple sources rhythmic signal, noise interference, and long-period features that exceed sampling window capabilities. Through experiments, our non-deep-learning synthetic data enables models to achieve superior reconstruction performance and universal explicit representation extraction without the need for real data.
format Preprint
id arxiv_https___arxiv_org_abs_2402_00607
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are Synthetic Time-series Data Really not as Good as Real Data?
Fu, Fanzhe
Chen, Junru
Zhang, Jing
Yang, Carl
Ma, Lvbin
Yang, Yang
Machine Learning
Artificial Intelligence
Time-series data presents limitations stemming from data quality issues, bias and vulnerabilities, and generalization problem. Integrating universal data synthesis methods holds promise in improving generalization. However, current methods cannot guarantee that the generator's output covers all unseen real data. In this paper, we introduce InfoBoost -- a highly versatile cross-domain data synthesizing framework with time series representation learning capability. We have developed a method based on synthetic data that enables model training without the need for real data, surpassing the performance of models trained with real data. Additionally, we have trained a universal feature extractor based on our synthetic data that is applicable to all time-series data. Our approach overcomes interference from multiple sources rhythmic signal, noise interference, and long-period features that exceed sampling window capabilities. Through experiments, our non-deep-learning synthetic data enables models to achieve superior reconstruction performance and universal explicit representation extraction without the need for real data.
title Are Synthetic Time-series Data Really not as Good as Real Data?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2402.00607