TabSynDex: A Universal Metric for Robust Evaluation of Synthetic Tabular Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chundawat, Vikram S, Tarun, Ayush K, Mandal, Murari, Lahoti, Mukund, Narang, Pratik
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911910193004544
author Chundawat, Vikram S
Tarun, Ayush K
Mandal, Murari
Lahoti, Mukund
Narang, Pratik
author_facet Chundawat, Vikram S
Tarun, Ayush K
Mandal, Murari
Lahoti, Mukund
Narang, Pratik
contents Synthetic tabular data generation becomes crucial when real data is limited, expensive to collect, or simply cannot be used due to privacy concerns. However, producing good quality synthetic data is challenging. Several probabilistic, statistical, generative adversarial networks (GANs), and variational auto-encoder (VAEs) based approaches have been presented for synthetic tabular data generation. Once generated, evaluating the quality of the synthetic data is quite challenging. Some of the traditional metrics have been used in the literature but there is lack of a common, robust, and single metric. This makes it difficult to properly compare the effectiveness of different synthetic tabular data generation methods. In this paper we propose a new universal metric, TabSynDex, for robust evaluation of synthetic data. The proposed metric assesses the similarity of synthetic data with real data through different component scores which evaluate the characteristics that are desirable for ``high quality'' synthetic data. Being a single score metric and having an implicit bound, TabSynDex can also be used to observe and evaluate the training of neural network based approaches. This would help in obtaining insights that was not possible earlier. We present several baseline models for comparative analysis of the proposed evaluation metric with existing generative models. We also give a comparative analysis between TabSynDex and existing synthetic tabular data evaluation metrics. This shows the effectiveness and universality of our metric over the existing metrics. Source Code: \url{https://github.com/vikram2000b/tabsyndex}
format Preprint
id arxiv_https___arxiv_org_abs_2207_05295
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle TabSynDex: A Universal Metric for Robust Evaluation of Synthetic Tabular Data
Chundawat, Vikram S
Tarun, Ayush K
Mandal, Murari
Lahoti, Mukund
Narang, Pratik
Machine Learning
Synthetic tabular data generation becomes crucial when real data is limited, expensive to collect, or simply cannot be used due to privacy concerns. However, producing good quality synthetic data is challenging. Several probabilistic, statistical, generative adversarial networks (GANs), and variational auto-encoder (VAEs) based approaches have been presented for synthetic tabular data generation. Once generated, evaluating the quality of the synthetic data is quite challenging. Some of the traditional metrics have been used in the literature but there is lack of a common, robust, and single metric. This makes it difficult to properly compare the effectiveness of different synthetic tabular data generation methods. In this paper we propose a new universal metric, TabSynDex, for robust evaluation of synthetic data. The proposed metric assesses the similarity of synthetic data with real data through different component scores which evaluate the characteristics that are desirable for ``high quality'' synthetic data. Being a single score metric and having an implicit bound, TabSynDex can also be used to observe and evaluate the training of neural network based approaches. This would help in obtaining insights that was not possible earlier. We present several baseline models for comparative analysis of the proposed evaluation metric with existing generative models. We also give a comparative analysis between TabSynDex and existing synthetic tabular data evaluation metrics. This shows the effectiveness and universality of our metric over the existing metrics. Source Code: \url{https://github.com/vikram2000b/tabsyndex}
title TabSynDex: A Universal Metric for Robust Evaluation of Synthetic Tabular Data
topic Machine Learning
url https://arxiv.org/abs/2207.05295