Saved in:
Bibliographic Details
Main Authors: Hsieh, Din-Yin, Wang, Chi-Hua, Cheng, Guang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2401.00965
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910285379403776
author Hsieh, Din-Yin
Wang, Chi-Hua
Cheng, Guang
author_facet Hsieh, Din-Yin
Wang, Chi-Hua
Cheng, Guang
contents Exploring generative model training for synthetic tabular data, specifically in sequential contexts such as credit card transaction data, presents significant challenges. This paper addresses these challenges, focusing on attaining both high fidelity to actual data and optimal utility for machine learning tasks. We introduce five pre-processing schemas to enhance the training of the Conditional Probabilistic Auto-Regressive Model (CPAR), demonstrating incremental improvements in the synthetic data's fidelity and utility. Upon achieving satisfactory fidelity levels, our attention shifts to training fraud detection models tailored for time-series data, evaluating the utility of the synthetic data. Our findings offer valuable insights and practical guidelines for synthetic data practitioners in the finance sector, transitioning from real to synthetic datasets for training purposes, and illuminating broader methodologies for synthesizing credit card transaction time series.
format Preprint
id arxiv_https___arxiv_org_abs_2401_00965
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improve Fidelity and Utility of Synthetic Credit Card Transaction Time Series from Data-centric Perspective
Hsieh, Din-Yin
Wang, Chi-Hua
Cheng, Guang
Machine Learning
Exploring generative model training for synthetic tabular data, specifically in sequential contexts such as credit card transaction data, presents significant challenges. This paper addresses these challenges, focusing on attaining both high fidelity to actual data and optimal utility for machine learning tasks. We introduce five pre-processing schemas to enhance the training of the Conditional Probabilistic Auto-Regressive Model (CPAR), demonstrating incremental improvements in the synthetic data's fidelity and utility. Upon achieving satisfactory fidelity levels, our attention shifts to training fraud detection models tailored for time-series data, evaluating the utility of the synthetic data. Our findings offer valuable insights and practical guidelines for synthetic data practitioners in the finance sector, transitioning from real to synthetic datasets for training purposes, and illuminating broader methodologies for synthesizing credit card transaction time series.
title Improve Fidelity and Utility of Synthetic Credit Card Transaction Time Series from Data-centric Perspective
topic Machine Learning
url https://arxiv.org/abs/2401.00965