Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gong, Chen, Liang, Bo, Gao, Wei, Xu, Chenren
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2506.23174
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908427636178944
author Gong, Chen
Liang, Bo
Gao, Wei
Xu, Chenren
author_facet Gong, Chen
Liang, Bo
Gao, Wei
Xu, Chenren
contents Generative models have gained significant attention for their ability to produce realistic synthetic data that supplements the quantity of real-world datasets. While recent studies show performance improvements in wireless sensing tasks by incorporating all synthetic data into training sets, the quality of synthetic data remains unpredictable and the resulting performance gains are not guaranteed. To address this gap, we propose tractable and generalizable metrics to quantify quality attributes of synthetic data - affinity and diversity. Our assessment reveals prevalent affinity limitation in current wireless synthetic data, leading to mislabeled data and degraded task performance. We attribute the quality limitation to generative models' lack of awareness of untrained conditions and domain-specific processing. To mitigate these issues, we introduce SynCheck, a quality-guided synthetic data utilization scheme that refines synthetic data quality during task model training. Our evaluation demonstrates that SynCheck consistently outperforms quality-oblivious utilization of synthetic data, and achieves 4.3% performance improvement even when the previous utilization degrades performance by 13.4%.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23174
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Can Speak for Itself: Quality-guided Utilization of Wireless Synthetic Data
Gong, Chen
Liang, Bo
Gao, Wei
Xu, Chenren
Machine Learning
Artificial Intelligence
Generative models have gained significant attention for their ability to produce realistic synthetic data that supplements the quantity of real-world datasets. While recent studies show performance improvements in wireless sensing tasks by incorporating all synthetic data into training sets, the quality of synthetic data remains unpredictable and the resulting performance gains are not guaranteed. To address this gap, we propose tractable and generalizable metrics to quantify quality attributes of synthetic data - affinity and diversity. Our assessment reveals prevalent affinity limitation in current wireless synthetic data, leading to mislabeled data and degraded task performance. We attribute the quality limitation to generative models' lack of awareness of untrained conditions and domain-specific processing. To mitigate these issues, we introduce SynCheck, a quality-guided synthetic data utilization scheme that refines synthetic data quality during task model training. Our evaluation demonstrates that SynCheck consistently outperforms quality-oblivious utilization of synthetic data, and achieves 4.3% performance improvement even when the previous utilization degrades performance by 13.4%.
title Data Can Speak for Itself: Quality-guided Utilization of Wireless Synthetic Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.23174