Evaluating Inter-Column Logical Relationships in Synthetic Tabular Data Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Long, Yunbo, Xu, Liming, Brintrup, Alexandra
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914572889227264
author Long, Yunbo
Xu, Liming
Brintrup, Alexandra
author_facet Long, Yunbo
Xu, Liming
Brintrup, Alexandra
contents Current evaluations of synthetic tabular data mainly focus on how well joint distributions are modeled, often overlooking the assessment of their effectiveness in preserving realistic event sequences and coherent entity relationships across columns.This paper proposes three evaluation metrics designed to assess the preservation of logical relationships among columns in synthetic tabular data. We validate these metrics by assessing the performance of both classical and state-of-the-art generation methods on a real-world industrial dataset.Experimental results reveal that existing methods often fail to rigorously maintain logical consistency (e.g., hierarchical relationships in geography or organization) and dependencies (e.g., temporal sequences or mathematical relationships), which are crucial for preserving the fine-grained realism of real-world tabular data. Building on these insights, this study also discusses possible pathways to better capture logical relationships while modeling the distribution of synthetic tabular data. The code is available at https://github.com/Yunbo-max/TabLogicEval.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Inter-Column Logical Relationships in Synthetic Tabular Data Generation
Long, Yunbo
Xu, Liming
Brintrup, Alexandra
Machine Learning
Current evaluations of synthetic tabular data mainly focus on how well joint distributions are modeled, often overlooking the assessment of their effectiveness in preserving realistic event sequences and coherent entity relationships across columns.This paper proposes three evaluation metrics designed to assess the preservation of logical relationships among columns in synthetic tabular data. We validate these metrics by assessing the performance of both classical and state-of-the-art generation methods on a real-world industrial dataset.Experimental results reveal that existing methods often fail to rigorously maintain logical consistency (e.g., hierarchical relationships in geography or organization) and dependencies (e.g., temporal sequences or mathematical relationships), which are crucial for preserving the fine-grained realism of real-world tabular data. Building on these insights, this study also discusses possible pathways to better capture logical relationships while modeling the distribution of synthetic tabular data. The code is available at https://github.com/Yunbo-max/TabLogicEval.
title Evaluating Inter-Column Logical Relationships in Synthetic Tabular Data Generation
topic Machine Learning
url https://arxiv.org/abs/2502.04055