When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models
Fuente:
arXiv
Saved in:
| Main Authors: | Mustaqim, S. M., Kotal, Anantaa, Yi, Paul H. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KIPPS: Knowledge infusion in Privacy Preserving Synthetic Data Generation
by: Kotal, Anantaa, et al.
Published: (2024)
by: Kotal, Anantaa, et al.
Published: (2024)
Differentially Private Synthetic Data Generation Using Context-Aware GANs
by: Kotal, Anantaa, et al.
Published: (2025)
by: Kotal, Anantaa, et al.
Published: (2025)
Impugan: Learning Conditional Generative Models for Robust Data Imputation
by: Mahmud, Zalish, et al.
Published: (2025)
by: Mahmud, Zalish, et al.
Published: (2025)
Privacy-Preserving Data Sharing in Agriculture: Enforcing Policy Rules for Secure and Confidential Data Synthesis
by: Kotal, Anantaa, et al.
Published: (2023)
by: Kotal, Anantaa, et al.
Published: (2023)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
by: Galeone, Cosimo, et al.
Published: (2026)
by: Galeone, Cosimo, et al.
Published: (2026)
Why Isn't Relational Learning Taking Over the World?
by: Poole, David
Published: (2025)
by: Poole, David
Published: (2025)
KiNETGAN: Enabling Distributed Network Intrusion Detection through Knowledge-Infused Synthetic Data Generation
by: Kotal, Anantaa, et al.
Published: (2024)
by: Kotal, Anantaa, et al.
Published: (2024)
FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning
by: Sanchez Jr., David, et al.
Published: (2025)
by: Sanchez Jr., David, et al.
Published: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
Generating High-quality Privacy-preserving Synthetic Data
by: Yavo, David, et al.
Published: (2026)
by: Yavo, David, et al.
Published: (2026)
Synthetic Data Privacy Metrics
by: Steier, Amy, et al.
Published: (2025)
by: Steier, Amy, et al.
Published: (2025)
Inverse Scaling: When Bigger Isn't Better
by: McKenzie, Ian R., et al.
Published: (2023)
by: McKenzie, Ian R., et al.
Published: (2023)
Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
PrivComp-KG : Leveraging Knowledge Graph and Large Language Models for Privacy Policy Compliance Verification
by: Garza, Leon, et al.
Published: (2024)
by: Garza, Leon, et al.
Published: (2024)
Hidden Leaks in Time Series Forecasting: How Data Leakage Affects LSTM Evaluation Across Configurations and Validation Strategies
by: Albelali, Salma, et al.
Published: (2025)
by: Albelali, Salma, et al.
Published: (2025)
Evaluating Privacy Leakage in Split Learning
by: Qiu, Xinchi, et al.
Published: (2023)
by: Qiu, Xinchi, et al.
Published: (2023)
Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
by: Freiesleben, Timo, et al.
Published: (2026)
by: Freiesleben, Timo, et al.
Published: (2026)
When Slower Isn't Truer: Inverse Scaling Law of Truthfulness in Multimodal Reasoning
by: Fang, Sitong, et al.
Published: (2025)
by: Fang, Sitong, et al.
Published: (2025)
When Alignment Isn't Enough: Response-Path Attacks on LLM Agents
by: Luo, Mingyu, et al.
Published: (2026)
by: Luo, Mingyu, et al.
Published: (2026)
Hidden-State Privacy Has an Empty Middle
by: Bell, Alexander Okezue
Published: (2026)
by: Bell, Alexander Okezue
Published: (2026)
How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
by: Ponomareva, Natalia, et al.
Published: (2025)
by: Ponomareva, Natalia, et al.
Published: (2025)
Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity
by: Eshbaugh, Jackson, et al.
Published: (2025)
by: Eshbaugh, Jackson, et al.
Published: (2025)
Retrieval-Augmented Framework for LLM-Based Clinical Decision Support
by: Garza, Leon, et al.
Published: (2025)
by: Garza, Leon, et al.
Published: (2025)
Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
by: Devanathan, Rishikesh, et al.
Published: (2025)
by: Devanathan, Rishikesh, et al.
Published: (2025)
Generating Synthetic Health Sensor Data for Privacy-Preserving Wearable Stress Detection
by: Lange, Lucas, et al.
Published: (2024)
by: Lange, Lucas, et al.
Published: (2024)
Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage
by: Rashid, Md Rafi Ur, et al.
Published: (2024)
by: Rashid, Md Rafi Ur, et al.
Published: (2024)
When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
by: Adamkiewicz, Krzysztof, et al.
Published: (2026)
by: Adamkiewicz, Krzysztof, et al.
Published: (2026)
A Note on Shumailov et al. (2024): `AI Models Collapse When Trained on Recursively Generated Data'
by: Borji, Ali
Published: (2024)
by: Borji, Ali
Published: (2024)
The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
by: Yao, Zexi, et al.
Published: (2025)
by: Yao, Zexi, et al.
Published: (2025)
SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data
by: Hu, Bing, et al.
Published: (2026)
by: Hu, Bing, et al.
Published: (2026)
Knowing Isn't Understanding: Re-grounding Generative Proactivity with Epistemic and Behavioral Insight
by: Kaur, Kirandeep, et al.
Published: (2026)
by: Kaur, Kirandeep, et al.
Published: (2026)
Synthetic Series-Symbol Data Generation for Time Series Foundation Models
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Generating Synthetic Net Load Data with Physics-informed Diffusion Model
by: Zhang, Shaorong, et al.
Published: (2024)
by: Zhang, Shaorong, et al.
Published: (2024)
Leakage and Interpretability in Concept-Based Models
by: Parisini, Enrico, et al.
Published: (2025)
by: Parisini, Enrico, et al.
Published: (2025)
ML Interpretability: Simple Isn't Easy
by: Räz, Tim
Published: (2022)
by: Räz, Tim
Published: (2022)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
Synthetic Data Generation for Augmenting Small Samples
by: Liu, Dan, et al.
Published: (2025)
by: Liu, Dan, et al.
Published: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
SurvDiff: A Diffusion Model for Generating Synthetic Data in Survival Analysis
by: Brockschmidt, Marie, et al.
Published: (2025)
by: Brockschmidt, Marie, et al.
Published: (2025)
Gradient Inversion Transcript: Leveraging Robust Generative Priors to Reconstruct Training Data from Gradient Leakage
by: Chen, Xinping, et al.
Published: (2025)
by: Chen, Xinping, et al.
Published: (2025)
Similar Items
-
KIPPS: Knowledge infusion in Privacy Preserving Synthetic Data Generation
by: Kotal, Anantaa, et al.
Published: (2024) -
Differentially Private Synthetic Data Generation Using Context-Aware GANs
by: Kotal, Anantaa, et al.
Published: (2025) -
Impugan: Learning Conditional Generative Models for Robust Data Imputation
by: Mahmud, Zalish, et al.
Published: (2025) -
Privacy-Preserving Data Sharing in Agriculture: Enforcing Policy Rules for Secure and Confidential Data Synthesis
by: Kotal, Anantaa, et al.
Published: (2023) -
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
by: Galeone, Cosimo, et al.
Published: (2026)