The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hallinan, Skyler, Jung, Jaehun, Sclar, Melanie, Lu, Ximing, Ravichander, Abhilasha, Ramnath, Sahana, Choi, Yejin, Karimireddy, Sai Praneeth, Mireshghallah, Niloofar, Ren, Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917435252146176
author Hallinan, Skyler
Jung, Jaehun
Sclar, Melanie
Lu, Ximing
Ravichander, Abhilasha
Ramnath, Sahana
Choi, Yejin
Karimireddy, Sai Praneeth
Mireshghallah, Niloofar
Ren, Xiang
author_facet Hallinan, Skyler
Jung, Jaehun
Sclar, Melanie
Lu, Ximing
Ravichander, Abhilasha
Ramnath, Sahana
Choi, Yejin
Karimireddy, Sai Praneeth
Mireshghallah, Niloofar
Ren, Xiang
contents Membership inference attacks serves as useful tool for fair use of language models, such as detecting potential copyright infringement and auditing data leakage. However, many current state-of-the-art attacks require access to models' hidden states or probability distribution, which prevents investigation into more widely-used, API-access only models like GPT-4. In this work, we introduce N-Gram Coverage Attack, a membership inference attack that relies solely on text outputs from the target model, enabling attacks on completely black-box models. We leverage the observation that models are more likely to memorize and subsequently generate text patterns that were commonly observed in their training data. Specifically, to make a prediction on a candidate member, N-Gram Coverage Attack first obtains multiple model generations conditioned on a prefix of the candidate. It then uses n-gram overlap metrics to compute and aggregate the similarities of these outputs with the ground truth suffix; high similarities indicate likely membership. We first demonstrate on a diverse set of existing benchmarks that N-Gram Coverage Attack outperforms other black-box methods while also impressively achieving comparable or even better performance to state-of-the-art white-box attacks - despite having access to only text outputs. Interestingly, we find that the success rate of our method scales with the attack compute budget - as we increase the number of sequences generated from the target model conditioned on the prefix, attack performance tends to improve. Having verified the accuracy of our method, we use it to investigate previously unstudied closed OpenAI models on multiple domains. We find that more recent models, such as GPT-4o, exhibit increased robustness to membership inference, suggesting an evolving trend toward improved privacy protections.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09603
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
Hallinan, Skyler
Jung, Jaehun
Sclar, Melanie
Lu, Ximing
Ravichander, Abhilasha
Ramnath, Sahana
Choi, Yejin
Karimireddy, Sai Praneeth
Mireshghallah, Niloofar
Ren, Xiang
Computation and Language
Membership inference attacks serves as useful tool for fair use of language models, such as detecting potential copyright infringement and auditing data leakage. However, many current state-of-the-art attacks require access to models' hidden states or probability distribution, which prevents investigation into more widely-used, API-access only models like GPT-4. In this work, we introduce N-Gram Coverage Attack, a membership inference attack that relies solely on text outputs from the target model, enabling attacks on completely black-box models. We leverage the observation that models are more likely to memorize and subsequently generate text patterns that were commonly observed in their training data. Specifically, to make a prediction on a candidate member, N-Gram Coverage Attack first obtains multiple model generations conditioned on a prefix of the candidate. It then uses n-gram overlap metrics to compute and aggregate the similarities of these outputs with the ground truth suffix; high similarities indicate likely membership. We first demonstrate on a diverse set of existing benchmarks that N-Gram Coverage Attack outperforms other black-box methods while also impressively achieving comparable or even better performance to state-of-the-art white-box attacks - despite having access to only text outputs. Interestingly, we find that the success rate of our method scales with the attack compute budget - as we increase the number of sequences generated from the target model conditioned on the prefix, attack performance tends to improve. Having verified the accuracy of our method, we use it to investigate previously unstudied closed OpenAI models on multiple domains. We find that more recent models, such as GPT-4o, exhibit increased robustness to membership inference, suggesting an evolving trend toward improved privacy protections.
title The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
topic Computation and Language
url https://arxiv.org/abs/2508.09603