LLM Dataset Inference: Did you train on my dataset?
Fuente:
arXiv
Saved in:
| Main Authors: | Maini, Pratyush, Jia, Hengrui, Papernot, Nicolas, Dziedzic, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
by: Rastogi, Saksham, et al.
Published: (2025)
by: Rastogi, Saksham, et al.
Published: (2025)
On the Privacy Risk of In-context Learning
by: Duan, Haonan, et al.
Published: (2024)
by: Duan, Haonan, et al.
Published: (2024)
Have it your way: Individualized Privacy Assignment for DP-SGD
by: Boenisch, Franziska, et al.
Published: (2023)
by: Boenisch, Franziska, et al.
Published: (2023)
Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
by: Meeus, Matthieu, et al.
Published: (2023)
by: Meeus, Matthieu, et al.
Published: (2023)
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
by: Thudi, Anvith, et al.
Published: (2023)
by: Thudi, Anvith, et al.
Published: (2023)
Decentralised, Collaborative, and Privacy-preserving Machine Learning for Multi-Hospital Data
by: Fang, Congyu, et al.
Published: (2024)
by: Fang, Congyu, et al.
Published: (2024)
Is poisoning a real threat to LLM alignment? Maybe more so than you think
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
Backdoor Detection through Replicated Execution of Outsourced Training
by: Jia, Hengrui, et al.
Published: (2025)
by: Jia, Hengrui, et al.
Published: (2025)
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
by: Park, Seong-Gyu, et al.
Published: (2026)
by: Park, Seong-Gyu, et al.
Published: (2026)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
by: Liu, Ken Ziyu, et al.
Published: (2025)
by: Liu, Ken Ziyu, et al.
Published: (2025)
Proving membership in LLM pretraining data via data watermarks
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
Tighter Privacy Auditing of DP-SGD in the Hidden State Threat Model
by: Cebere, Tudor, et al.
Published: (2024)
by: Cebere, Tudor, et al.
Published: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
ADAGE: Active Defenses Against GNN Extraction
by: Xu, Jing, et al.
Published: (2025)
by: Xu, Jing, et al.
Published: (2025)
Context-Aware Membership Inference Attacks against Pre-trained Large Language Models
by: Chang, Hongyan, et al.
Published: (2024)
by: Chang, Hongyan, et al.
Published: (2024)
Membership Inference Attacks and Privacy in Topic Modeling
by: Manzonelli, Nico, et al.
Published: (2024)
by: Manzonelli, Nico, et al.
Published: (2024)
User Inference Attacks on Large Language Models
by: Kandpal, Nikhil, et al.
Published: (2023)
by: Kandpal, Nikhil, et al.
Published: (2023)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
by: Fu, Wenjie, et al.
Published: (2024)
by: Fu, Wenjie, et al.
Published: (2024)
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
by: Zhang, Anqi, et al.
Published: (2024)
by: Zhang, Anqi, et al.
Published: (2024)
Blind Baselines Beat Membership Inference Attacks for Foundation Models
by: Das, Debeshee, et al.
Published: (2024)
by: Das, Debeshee, et al.
Published: (2024)
VERA: Variational Inference Framework for Jailbreaking Large Language Models
by: Lochab, Anamika, et al.
Published: (2025)
by: Lochab, Anamika, et al.
Published: (2025)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
by: Frikha, Ahmed, et al.
Published: (2024)
by: Frikha, Ahmed, et al.
Published: (2024)
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
by: Meeus, Matthieu, et al.
Published: (2024)
by: Meeus, Matthieu, et al.
Published: (2024)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
by: Sahili, Ali Al, et al.
Published: (2025)
by: Sahili, Ali Al, et al.
Published: (2025)
LLM Unlearning Should Be Form-Independent
by: Ye, Xiaotian, et al.
Published: (2025)
by: Ye, Xiaotian, et al.
Published: (2025)
GCG Attack On A Diffusion LLM
by: Neyroud, Ruben, et al.
Published: (2025)
by: Neyroud, Ruben, et al.
Published: (2025)
SecFormer: Fast and Accurate Privacy-Preserving Inference for Transformer Models via SMPC
by: Luo, Jinglong, et al.
Published: (2024)
by: Luo, Jinglong, et al.
Published: (2024)
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
by: Nguyen, Khanh, et al.
Published: (2025)
by: Nguyen, Khanh, et al.
Published: (2025)
LLMGuard: Guarding Against Unsafe LLM Behavior
by: Goyal, Shubh, et al.
Published: (2024)
by: Goyal, Shubh, et al.
Published: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
by: Assogba, Yannick, et al.
Published: (2026)
by: Assogba, Yannick, et al.
Published: (2026)
Localizing Malicious Outputs from CodeLLM
by: Borana, Mayukh, et al.
Published: (2025)
by: Borana, Mayukh, et al.
Published: (2025)
The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
by: Makroo, Owais, et al.
Published: (2025)
by: Makroo, Owais, et al.
Published: (2025)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
by: Liu, Bing, et al.
Published: (2026)
by: Liu, Bing, et al.
Published: (2026)
CDI: Copyrighted Data Identification in Diffusion Models
by: Dubiński, Jan, et al.
Published: (2024)
by: Dubiński, Jan, et al.
Published: (2024)
Similar Items
-
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
by: Rastogi, Saksham, et al.
Published: (2025) -
On the Privacy Risk of In-context Learning
by: Duan, Haonan, et al.
Published: (2024) -
Have it your way: Individualized Privacy Assignment for DP-SGD
by: Boenisch, Franziska, et al.
Published: (2023) -
Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
by: Meeus, Matthieu, et al.
Published: (2023) -
Gradients Look Alike: Sensitivity is Often Overestimated in DP-SGD
by: Thudi, Anvith, et al.
Published: (2023)