Exploring the limits of strong membership inference attacks on large language models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hayes, Jamie, Shumailov, Ilia, Choquette-Choo, Christopher A., Jagielski, Matthew, Kaissis, George, Nasr, Milad, Ghalebikesabi, Sahra, Annamalai, Meenatchi Sundaram Mutu Selva, Mireshghallah, Niloofar, Shilov, Igor, Meeus, Matthieu, de Montjoye, Yves-Alexandre, Lee, Katherine, Boenisch, Franziska, Dziedzic, Adam, Cooper, A. Feder
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917188747657216
author Hayes, Jamie
Shumailov, Ilia
Choquette-Choo, Christopher A.
Jagielski, Matthew
Kaissis, George
Nasr, Milad
Ghalebikesabi, Sahra
Annamalai, Meenatchi Sundaram Mutu Selva
Mireshghallah, Niloofar
Shilov, Igor
Meeus, Matthieu
de Montjoye, Yves-Alexandre
Lee, Katherine
Boenisch, Franziska
Dziedzic, Adam
Cooper, A. Feder
author_facet Hayes, Jamie
Shumailov, Ilia
Choquette-Choo, Christopher A.
Jagielski, Matthew
Kaissis, George
Nasr, Milad
Ghalebikesabi, Sahra
Annamalai, Meenatchi Sundaram Mutu Selva
Mireshghallah, Niloofar
Shilov, Igor
Meeus, Matthieu
de Montjoye, Yves-Alexandre
Lee, Katherine
Boenisch, Franziska
Dziedzic, Adam
Cooper, A. Feder
contents State-of-the-art membership inference attacks (MIAs) typically require training many reference models, making it difficult to scale these attacks to large pre-trained language models (LLMs). As a result, prior research has either relied on weaker attacks that avoid training references (e.g., fine-tuning attacks), or on stronger attacks applied to small models and datasets. However, weaker attacks have been shown to be brittle and insights from strong attacks in simplified settings do not translate to today's LLMs. These challenges prompt an important question: are the limitations observed in prior work due to attack design choices, or are MIAs fundamentally ineffective on LLMs? We address this question by scaling LiRA--one of the strongest MIAs--to GPT-2 architectures ranging from 10M to 1B parameters, training references on over 20B tokens from the C4 dataset. Our results advance the understanding of MIAs on LLMs in four key ways. While (1) strong MIAs can succeed on pre-trained LLMs, (2) their effectiveness, remains limited (e.g., AUC<0.7) in practical settings. (3) Even when strong MIAs achieve better-than-random AUC, aggregate metrics can conceal substantial per-sample MIA decision instability: due to training randomness, many decisions are so unstable that they are statistically indistinguishable from a coin flip. Finally, (4) the relationship between MIA success and related LLM privacy metrics is not as straightforward as prior work has suggested.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the limits of strong membership inference attacks on large language models
Hayes, Jamie
Shumailov, Ilia
Choquette-Choo, Christopher A.
Jagielski, Matthew
Kaissis, George
Nasr, Milad
Ghalebikesabi, Sahra
Annamalai, Meenatchi Sundaram Mutu Selva
Mireshghallah, Niloofar
Shilov, Igor
Meeus, Matthieu
de Montjoye, Yves-Alexandre
Lee, Katherine
Boenisch, Franziska
Dziedzic, Adam
Cooper, A. Feder
Cryptography and Security
Artificial Intelligence
Machine Learning
State-of-the-art membership inference attacks (MIAs) typically require training many reference models, making it difficult to scale these attacks to large pre-trained language models (LLMs). As a result, prior research has either relied on weaker attacks that avoid training references (e.g., fine-tuning attacks), or on stronger attacks applied to small models and datasets. However, weaker attacks have been shown to be brittle and insights from strong attacks in simplified settings do not translate to today's LLMs. These challenges prompt an important question: are the limitations observed in prior work due to attack design choices, or are MIAs fundamentally ineffective on LLMs? We address this question by scaling LiRA--one of the strongest MIAs--to GPT-2 architectures ranging from 10M to 1B parameters, training references on over 20B tokens from the C4 dataset. Our results advance the understanding of MIAs on LLMs in four key ways. While (1) strong MIAs can succeed on pre-trained LLMs, (2) their effectiveness, remains limited (e.g., AUC<0.7) in practical settings. (3) Even when strong MIAs achieve better-than-random AUC, aggregate metrics can conceal substantial per-sample MIA decision instability: due to training randomness, many decisions are so unstable that they are statistically indistinguishable from a coin flip. Finally, (4) the relationship between MIA success and related LLM privacy metrics is not as straightforward as prior work has suggested.
title Exploring the limits of strong membership inference attacks on large language models
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.18773