Membership Inference Attacks on Tokenizers of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tong, Meng, Du, Yuntao, Chen, Kejiang, Zhang, Weiming, Li, Ninghui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913160801288192
author Tong, Meng
Du, Yuntao
Chen, Kejiang
Zhang, Weiming
Li, Ninghui
author_facet Tong, Meng
Du, Yuntao
Chen, Kejiang
Zhang, Weiming
Li, Ninghui
contents Membership inference attacks (MIAs) are widely used to assess the privacy risks associated with machine learning models. However, when these attacks are applied to pre-trained large language models (LLMs), they encounter significant challenges, including mislabeled samples, distribution shifts, and discrepancies in model size between experimental and real-world settings. To address these limitations, we introduce tokenizers as a new attack vector for membership inference. Specifically, a tokenizer converts raw text into tokens for LLMs. Unlike full models, tokenizers can be efficiently trained from scratch, thereby avoiding the aforementioned challenges. In addition, the tokenizer's training data is typically representative of the data used to pre-train LLMs. Despite these advantages, the potential of tokenizers as an attack vector remains unexplored. To this end, we present the first study on membership leakage through tokenizers and explore five attack methods to infer dataset membership. Extensive experiments on millions of Internet samples reveal the vulnerabilities in the tokenizers of state-of-the-art LLMs. To mitigate this emerging risk, we further propose an adaptive defense. Our findings highlight tokenizers as an overlooked yet critical privacy threat, underscoring the urgent need for privacy-preserving mechanisms specifically designed for them.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05699
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Membership Inference Attacks on Tokenizers of Large Language Models
Tong, Meng
Du, Yuntao
Chen, Kejiang
Zhang, Weiming
Li, Ninghui
Cryptography and Security
Artificial Intelligence
Membership inference attacks (MIAs) are widely used to assess the privacy risks associated with machine learning models. However, when these attacks are applied to pre-trained large language models (LLMs), they encounter significant challenges, including mislabeled samples, distribution shifts, and discrepancies in model size between experimental and real-world settings. To address these limitations, we introduce tokenizers as a new attack vector for membership inference. Specifically, a tokenizer converts raw text into tokens for LLMs. Unlike full models, tokenizers can be efficiently trained from scratch, thereby avoiding the aforementioned challenges. In addition, the tokenizer's training data is typically representative of the data used to pre-train LLMs. Despite these advantages, the potential of tokenizers as an attack vector remains unexplored. To this end, we present the first study on membership leakage through tokenizers and explore five attack methods to infer dataset membership. Extensive experiments on millions of Internet samples reveal the vulnerabilities in the tokenizers of state-of-the-art LLMs. To mitigate this emerging risk, we further propose an adaptive defense. Our findings highlight tokenizers as an overlooked yet critical privacy threat, underscoring the urgent need for privacy-preserving mechanisms specifically designed for them.
title Membership Inference Attacks on Tokenizers of Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.05699