Membership Inference Attacks on Sequence Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rossi, Lorenzo, Aerni, Michael, Zhang, Jie, Tramèr, Florian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912415604539392
author Rossi, Lorenzo
Aerni, Michael
Zhang, Jie
Tramèr, Florian
author_facet Rossi, Lorenzo
Aerni, Michael
Zhang, Jie
Tramèr, Florian
contents Sequence models, such as Large Language Models (LLMs) and autoregressive image generators, have a tendency to memorize and inadvertently leak sensitive information. While this tendency has critical legal implications, existing tools are insufficient to audit the resulting risks. We hypothesize that those tools' shortcomings are due to mismatched assumptions. Thus, we argue that effectively measuring privacy leakage in sequence models requires leveraging the correlations inherent in sequential generation. To illustrate this, we adapt a state-of-the-art membership inference attack to explicitly model within-sequence correlations, thereby demonstrating how a strong existing attack can be naturally extended to suit the structure of sequence models. Through a case study, we show that our adaptations consistently improve the effectiveness of memorization audits without introducing additional computational costs. Our work hence serves as an important stepping stone toward reliable memorization audits for large sequence models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05126
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Membership Inference Attacks on Sequence Models
Rossi, Lorenzo
Aerni, Michael
Zhang, Jie
Tramèr, Florian
Cryptography and Security
Machine Learning
Sequence models, such as Large Language Models (LLMs) and autoregressive image generators, have a tendency to memorize and inadvertently leak sensitive information. While this tendency has critical legal implications, existing tools are insufficient to audit the resulting risks. We hypothesize that those tools' shortcomings are due to mismatched assumptions. Thus, we argue that effectively measuring privacy leakage in sequence models requires leveraging the correlations inherent in sequential generation. To illustrate this, we adapt a state-of-the-art membership inference attack to explicitly model within-sequence correlations, thereby demonstrating how a strong existing attack can be naturally extended to suit the structure of sequence models. Through a case study, we show that our adaptations consistently improve the effectiveness of memorization audits without introducing additional computational costs. Our work hence serves as an important stepping stone toward reliable memorization audits for large sequence models.
title Membership Inference Attacks on Sequence Models
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2506.05126