Exploring the limits of strong membership inference attacks on large language models
Fuente:
arXiv
Saved in:
| Main Authors: | Hayes, Jamie, Shumailov, Ilia, Choquette-Choo, Christopher A., Jagielski, Matthew, Kaissis, George, Nasr, Milad, Ghalebikesabi, Sahra, Annamalai, Meenatchi Sundaram Mutu Selva, Mireshghallah, Niloofar, Shilov, Igor, Meeus, Matthieu, de Montjoye, Yves-Alexandre, Lee, Katherine, Boenisch, Franziska, Dziedzic, Adam, Cooper, A. Feder |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
It's Our Loss: No Privacy Amplification for Hidden State DP-SGD With Non-Convex Loss
by: Annamalai, Meenatchi Sundaram Muthu Selva
Published: (2024)
by: Annamalai, Meenatchi Sundaram Muthu Selva
Published: (2024)
Counterfactual Influence as a Distributional Quantity
by: Meeus, Matthieu, et al.
Published: (2025)
by: Meeus, Matthieu, et al.
Published: (2025)
Nearly Tight Black-Box Auditing of Differentially Private Machine Learning
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024)
The Hitchhiker's Guide to Efficient, End-to-End, and Tight DP Auditing
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2025)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2025)
Auditing Private Prediction
by: Chadha, Karan, et al.
Published: (2024)
by: Chadha, Karan, et al.
Published: (2024)
Tight Auditing of Differential Privacy in MST and AIM
by: Ganev, Georgi, et al.
Published: (2026)
by: Ganev, Georgi, et al.
Published: (2026)
A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic Data
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2023)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2023)
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
by: Borkar, Jaydeep, et al.
Published: (2025)
by: Borkar, Jaydeep, et al.
Published: (2025)
Measuring memorization in language models via probabilistic extraction
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
The Elusive Pursuit of Reproducing PATE-GAN: Benchmarking, Auditing, Debugging
by: Ganev, Georgi, et al.
Published: (2024)
by: Ganev, Georgi, et al.
Published: (2024)
CLIOPATRA: Extracting Private Information from LLM Insights
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2026)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2026)
"What do you want from theory alone?" Experimenting with Tight Auditing of Differentially Private Synthetic Data Generation
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024)
Beyond the Crawl: Unmasking Browser Fingerprinting in Real User Interactions
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2025)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2025)
Synthetic is all you need: removing the auxiliary data assumption for membership inference attacks against synthetic data
by: Guépin, Florent, et al.
Published: (2023)
by: Guépin, Florent, et al.
Published: (2023)
The Mosaic Memory of Large Language Models
by: Shilov, Igor, et al.
Published: (2024)
by: Shilov, Igor, et al.
Published: (2024)
Curation Leaks: Membership Inference Attacks against Data Curation for Machine Learning
by: Wahdany, Dariush, et al.
Published: (2026)
by: Wahdany, Dariush, et al.
Published: (2026)
Differentially Private Prototypes for Imbalanced Transfer Learning
by: Wahdany, Dariush, et al.
Published: (2024)
by: Wahdany, Dariush, et al.
Published: (2024)
To Shuffle or not to Shuffle: Auditing DP-SGD with Shuffling
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024)
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024)
Understanding the Impact of Data Domain Extraction on Synthetic Data Privacy
by: Ganev, Georgi, et al.
Published: (2025)
by: Ganev, Georgi, et al.
Published: (2025)
The Importance of Being Discrete: Measuring the Impact of Discretization in End-to-End Differentially Private Synthetic Data
by: Ganev, Georgi, et al.
Published: (2025)
by: Ganev, Georgi, et al.
Published: (2025)
A Unified Framework for Adversary-Aware Differential Privacy Bounds
by: Swanberg, Marika, et al.
Published: (2025)
by: Swanberg, Marika, et al.
Published: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs
by: Wang, Xun, et al.
Published: (2025)
by: Wang, Xun, et al.
Published: (2025)
Copyright Traps for Large Language Models
by: Meeus, Matthieu, et al.
Published: (2024)
by: Meeus, Matthieu, et al.
Published: (2024)
Extracting alignment data in open models
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
LLMs unlock new paths to monetizing exploits
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
by: Dahiya, Pranav, et al.
Published: (2023)
by: Dahiya, Pranav, et al.
Published: (2023)
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
by: Meeus, Matthieu, et al.
Published: (2024)
by: Meeus, Matthieu, et al.
Published: (2024)
Localizing and Mitigating Memorization in Image Autoregressive Models
by: Kasliwal, Aditya, et al.
Published: (2025)
by: Kasliwal, Aditya, et al.
Published: (2025)
Implementing Adaptations for Vision AutoRegressive Model
by: Shaikh, Kaif, et al.
Published: (2025)
by: Shaikh, Kaif, et al.
Published: (2025)
ADAGE: Active Defenses Against GNN Extraction
by: Xu, Jing, et al.
Published: (2025)
by: Xu, Jing, et al.
Published: (2025)
Position: Privacy Is Not Just Memorization!
by: Mireshghallah, Niloofar, et al.
Published: (2025)
by: Mireshghallah, Niloofar, et al.
Published: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
by: Steinke, Thomas, et al.
Published: (2024)
by: Steinke, Thomas, et al.
Published: (2024)
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
Similar Items
-
It's Our Loss: No Privacy Amplification for Hidden State DP-SGD With Non-Convex Loss
by: Annamalai, Meenatchi Sundaram Muthu Selva
Published: (2024) -
Counterfactual Influence as a Distributional Quantity
by: Meeus, Matthieu, et al.
Published: (2025) -
Nearly Tight Black-Box Auditing of Differentially Private Machine Learning
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2024) -
The Hitchhiker's Guide to Efficient, End-to-End, and Tight DP Auditing
by: Annamalai, Meenatchi Sundaram Muthu Selva, et al.
Published: (2025) -
Auditing Private Prediction
by: Chadha, Karan, et al.
Published: (2024)