Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yixin, Yang, Ziqing, Shen, Yun, Backes, Michael, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Challenge of Identifying the Origin of Black-Box Large Language Models
by: Yang, Ziqing, et al.
Published: (2025)
by: Yang, Ziqing, et al.
Published: (2025)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective
by: Ganev, Georgi, et al.
Published: (2026)
by: Ganev, Georgi, et al.
Published: (2026)
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
by: Yang, Ziqing, et al.
Published: (2024)
by: Yang, Ziqing, et al.
Published: (2024)
Voice Jailbreak Attacks Against GPT-4o
by: Shen, Xinyue, et al.
Published: (2024)
by: Shen, Xinyue, et al.
Published: (2024)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
by: Meeus, Matthieu, et al.
Published: (2025)
by: Meeus, Matthieu, et al.
Published: (2025)
Sharing is CAIRing: Characterizing Principles and Assessing Properties of Universal Privacy Evaluation for Synthetic Tabular Data
by: Hyrup, Tobias, et al.
Published: (2023)
by: Hyrup, Tobias, et al.
Published: (2023)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
Trustless Audits without Revealing Data or Models
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
by: Waiwitlikhit, Suppakit, et al.
Published: (2024)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
Gaming the Metric, Not the Harm: Certifying Safety Audits against Strategic Platform Manipulation
by: Burnat, Florian A. D., et al.
Published: (2026)
by: Burnat, Florian A. D., et al.
Published: (2026)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
by: Wen, Rui, et al.
Published: (2024)
by: Wen, Rui, et al.
Published: (2024)
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Watermarking LLM-Generated Datasets in Downstream Tasks
by: Liu, Yugeng, et al.
Published: (2025)
by: Liu, Yugeng, et al.
Published: (2025)
MGTBench: Benchmarking Machine-Generated Text Detection
by: He, Xinlei, et al.
Published: (2023)
by: He, Xinlei, et al.
Published: (2023)
Machine Unlearning: Taxonomy, Metrics, Applications, Challenges, and Prospects
by: Li, Na, et al.
Published: (2024)
by: Li, Na, et al.
Published: (2024)
End to End Collaborative Synthetic Data Generation
by: Pentyala, Sikha, et al.
Published: (2024)
by: Pentyala, Sikha, et al.
Published: (2024)
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
by: Jiang, Yukun, et al.
Published: (2024)
by: Jiang, Yukun, et al.
Published: (2024)
Synthetic Data: Revisiting the Privacy-Utility Trade-off
by: Sarmin, Fatima Jahan, et al.
Published: (2024)
by: Sarmin, Fatima Jahan, et al.
Published: (2024)
The Data Sharing Paradox of Synthetic Data in Healthcare
by: Achterberg, Jim, et al.
Published: (2025)
by: Achterberg, Jim, et al.
Published: (2025)
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
by: Wen, Rui, et al.
Published: (2025)
by: Wen, Rui, et al.
Published: (2025)
Peering Behind the Shield: Guardrail Identification in Large Language Models
by: Yang, Ziqing, et al.
Published: (2025)
by: Yang, Ziqing, et al.
Published: (2025)
Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
by: Akkus, Atilla, et al.
Published: (2024)
by: Akkus, Atilla, et al.
Published: (2024)
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
by: Bicakci, Kemal
Published: (2026)
by: Bicakci, Kemal
Published: (2026)
Link Stealing Attacks Against Inductive Graph Neural Networks
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric
by: Shvartzshnaider, Yan, et al.
Published: (2024)
by: Shvartzshnaider, Yan, et al.
Published: (2024)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
by: Wu, Yixin, et al.
Published: (2023)
by: Wu, Yixin, et al.
Published: (2023)
Scalable and Privacy-Preserving Synthetic Data Generation on Decentralised Web
by: Ramesh, Vishal, et al.
Published: (2023)
by: Ramesh, Vishal, et al.
Published: (2023)
FLAIM: AIM-based Synthetic Data Generation in the Federated Setting
by: Maddock, Samuel, et al.
Published: (2023)
by: Maddock, Samuel, et al.
Published: (2023)
Rethinking Data Protection in the (Generative) Artificial Intelligence Era
by: Li, Yiming, et al.
Published: (2025)
by: Li, Yiming, et al.
Published: (2025)
Synthetic Data and Health Privacy
by: Abgrall, Gwénolé, et al.
Published: (2025)
by: Abgrall, Gwénolé, et al.
Published: (2025)
Does Training with Synthetic Data Truly Protect Privacy?
by: Zhao, Yunpeng, et al.
Published: (2025)
by: Zhao, Yunpeng, et al.
Published: (2025)
Conscious Data Contribution via Community-Driven Chain-of-Thought Distillation
by: Libon, Lena, et al.
Published: (2025)
by: Libon, Lena, et al.
Published: (2025)
Similar Items
-
The Challenge of Identifying the Origin of Black-Box Large Language Models
by: Yang, Ziqing, et al.
Published: (2025) -
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
by: Wu, Yixin, et al.
Published: (2024) -
Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective
by: Ganev, Georgi, et al.
Published: (2026) -
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025) -
SOS! Soft Prompt Attack Against Open-Source Large Language Models
by: Yang, Ziqing, et al.
Published: (2024)