When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ward, Joshua, Gu, Bochao, Wang, Chi-Hua, Cheng, Guang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ensembling Membership Inference Attacks Against Tabular Generative Models
von: Ward, Joshua, et al.
Veröffentlicht: (2025)
von: Ward, Joshua, et al.
Veröffentlicht: (2025)
Data Plagiarism Index: Characterizing the Privacy Risk of Data-Copying in Tabular Generative Models
von: Ward, Joshua, et al.
Veröffentlicht: (2024)
von: Ward, Joshua, et al.
Veröffentlicht: (2024)
Finding Connections: Membership Inference Attacks for the Multi-Table Synthetic Data Setting
von: Ward, Joshua, et al.
Veröffentlicht: (2026)
von: Ward, Joshua, et al.
Veröffentlicht: (2026)
The Pitfalls of Memorization: When Memorization Hurts Generalization
von: Bayat, Reza, et al.
Veröffentlicht: (2024)
von: Bayat, Reza, et al.
Veröffentlicht: (2024)
Privacy Auditing Synthetic Data Release through Local Likelihood Attacks
von: Ward, Joshua, et al.
Veröffentlicht: (2025)
von: Ward, Joshua, et al.
Veröffentlicht: (2025)
Synth-MIA: A Testbed for Auditing Privacy Leakage in Tabular Data Synthesis
von: Ward, Joshua, et al.
Veröffentlicht: (2025)
von: Ward, Joshua, et al.
Veröffentlicht: (2025)
Downstream Task-Oriented Generative Model Selections on Synthetic Data Training for Fraud Detection Models
von: Cheng, Yinan, et al.
Veröffentlicht: (2024)
von: Cheng, Yinan, et al.
Veröffentlicht: (2024)
RePCS: Diagnosing Data Memorization in LLM-Powered Retrieval-Augmented Generation
von: Anh, Le Vu, et al.
Veröffentlicht: (2025)
von: Anh, Le Vu, et al.
Veröffentlicht: (2025)
TabAttackBench: A Benchmark for Adversarial Attacks on Tabular Data
von: He, Zhipeng, et al.
Veröffentlicht: (2025)
von: He, Zhipeng, et al.
Veröffentlicht: (2025)
TAGAL: Tabular Data Generation using Agentic LLM Methods
von: Ronval, Benoît, et al.
Veröffentlicht: (2025)
von: Ronval, Benoît, et al.
Veröffentlicht: (2025)
Memorization Sinks: Isolating Memorization during LLM Training
von: Ghosal, Gaurav R., et al.
Veröffentlicht: (2025)
von: Ghosal, Gaurav R., et al.
Veröffentlicht: (2025)
Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
Crafting Imperceptible On-Manifold Adversarial Attacks for Tabular Data
von: He, Zhipeng, et al.
Veröffentlicht: (2025)
von: He, Zhipeng, et al.
Veröffentlicht: (2025)
LLM Embeddings for Deep Learning on Tabular Data
von: Koloski, Boshko, et al.
Veröffentlicht: (2025)
von: Koloski, Boshko, et al.
Veröffentlicht: (2025)
When Do Neural Nets Outperform Boosted Trees on Tabular Data?
von: McElfresh, Duncan, et al.
Veröffentlicht: (2023)
von: McElfresh, Duncan, et al.
Veröffentlicht: (2023)
Membership and Memorization in LLM Knowledge Distillation
von: Zhang, Ziqi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2025)
Team, Then Trim: An Assembly-Line LLM Framework for High-Quality Tabular Data Generation
von: Zhang, Congjing, et al.
Veröffentlicht: (2026)
von: Zhang, Congjing, et al.
Veröffentlicht: (2026)
Watermarking Generative Categorical Data
von: Gu, Bochao, et al.
Veröffentlicht: (2024)
von: Gu, Bochao, et al.
Veröffentlicht: (2024)
TimeAutoDiff: A Unified Framework for Generation, Imputation, Forecasting, and Time-Varying Metadata Conditioning of Heterogeneous Time Series Tabular Data
von: Suh, Namjoon, et al.
Veröffentlicht: (2024)
von: Suh, Namjoon, et al.
Veröffentlicht: (2024)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
von: Barron, Joshua, et al.
Veröffentlicht: (2025)
von: Barron, Joshua, et al.
Veröffentlicht: (2025)
Measuring LLM Sensitivity in Transformer-based Tabular Data Synthesis
von: R, Maria F. Davila, et al.
Veröffentlicht: (2025)
von: R, Maria F. Davila, et al.
Veröffentlicht: (2025)
Automatic Demonstration Selection for LLM-based Tabular Data Classification
von: Han, Shuchu, et al.
Veröffentlicht: (2025)
von: Han, Shuchu, et al.
Veröffentlicht: (2025)
Doubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphs
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
Tabular Data Generation using Binary Diffusion
von: Kinakh, Vitaliy, et al.
Veröffentlicht: (2024)
von: Kinakh, Vitaliy, et al.
Veröffentlicht: (2024)
Risk In Context: Benchmarking Privacy Leakage of Foundation Models in Synthetic Tabular Data Generation
von: Byun, Jessup, et al.
Veröffentlicht: (2025)
von: Byun, Jessup, et al.
Veröffentlicht: (2025)
Critical Windows of Complexity Control: When Transformers Decide to Reason or Memorize
von: Ali, Sarwan
Veröffentlicht: (2026)
von: Ali, Sarwan
Veröffentlicht: (2026)
Diffusion-Driven Synthetic Tabular Data Generation for Enhanced DoS/DDoS Attack Classification
von: B, Aravind, et al.
Veröffentlicht: (2026)
von: B, Aravind, et al.
Veröffentlicht: (2026)
Batch Normalization Amplifies Memorization and Privacy Risks
von: Doan, Ngoc Phu, et al.
Veröffentlicht: (2026)
von: Doan, Ngoc Phu, et al.
Veröffentlicht: (2026)
Improving LLM Group Fairness on Tabular Data via In-Context Learning
von: Cherepanova, Valeriia, et al.
Veröffentlicht: (2024)
von: Cherepanova, Valeriia, et al.
Veröffentlicht: (2024)
Investigating Imperceptibility of Adversarial Attacks on Tabular Data: An Empirical Analysis
von: He, Zhipeng, et al.
Veröffentlicht: (2024)
von: He, Zhipeng, et al.
Veröffentlicht: (2024)
GEM-T: Generative Tabular Data via Fitting Moments
von: Li, Miao, et al.
Veröffentlicht: (2025)
von: Li, Miao, et al.
Veröffentlicht: (2025)
Diffusion Transformers for Tabular Data Time Series Generation
von: Garuti, Fabrizio, et al.
Veröffentlicht: (2025)
von: Garuti, Fabrizio, et al.
Veröffentlicht: (2025)
Generating Realistic Tabular Data with Large Language Models
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
TabularQGAN: A Quantum Generative Model for Tabular Data
von: Bhardwaj, Pallavi, et al.
Veröffentlicht: (2025)
von: Bhardwaj, Pallavi, et al.
Veröffentlicht: (2025)
How Do Flow Matching Models Memorize and Generalize in Sample Data Subspaces?
von: Gao, Weiguo, et al.
Veröffentlicht: (2024)
von: Gao, Weiguo, et al.
Veröffentlicht: (2024)
Exploring Transformer Placement in Variational Autoencoders for Tabular Data Generation
von: Silva, Aníbal, et al.
Veröffentlicht: (2026)
von: Silva, Aníbal, et al.
Veröffentlicht: (2026)
Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes
von: Jiang, Mingxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Mingxuan, et al.
Veröffentlicht: (2025)
TableGPT2: A Large Multimodal Model with Tabular Data Integration
von: Su, Aofeng, et al.
Veröffentlicht: (2024)
von: Su, Aofeng, et al.
Veröffentlicht: (2024)
PLeak: Prompt Leaking Attacks against Large Language Model Applications
von: Hui, Bo, et al.
Veröffentlicht: (2024)
von: Hui, Bo, et al.
Veröffentlicht: (2024)
Beyond Frequency: The Role of Redundancy in Large Language Model Memorization
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ensembling Membership Inference Attacks Against Tabular Generative Models
von: Ward, Joshua, et al.
Veröffentlicht: (2025) -
Data Plagiarism Index: Characterizing the Privacy Risk of Data-Copying in Tabular Generative Models
von: Ward, Joshua, et al.
Veröffentlicht: (2024) -
Finding Connections: Membership Inference Attacks for the Multi-Table Synthetic Data Setting
von: Ward, Joshua, et al.
Veröffentlicht: (2026) -
The Pitfalls of Memorization: When Memorization Hurts Generalization
von: Bayat, Reza, et al.
Veröffentlicht: (2024) -
Privacy Auditing Synthetic Data Release through Local Likelihood Attacks
von: Ward, Joshua, et al.
Veröffentlicht: (2025)