Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zeyuan, Ma, Yihan, Shen, Xinyue, Backes, Michael, Zhang, Yang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915989610823680
author Chen, Zeyuan
Ma, Yihan
Shen, Xinyue
Backes, Michael
Zhang, Yang
author_facet Chen, Zeyuan
Ma, Yihan
Shen, Xinyue
Backes, Michael
Zhang, Yang
contents Large language models (LLMs) show strong performance across many applications, but their ability to memorize and potentially reveal training data raises serious privacy concerns. We introduce the PopQuiz Attack, a black-box membership inference attack that tests whether a model can recall specific training examples. The core idea is to turn target data into quiz-style multiple-choice questions and infer membership from the model's answers. Across six widely used LLMs (GPT-3.5, GPT-4o, LLaMA2-7b, LLaMA2-13b, Mistral-7b, and Vicuna-7b) and four datasets, our method achieves an average ROC-AUC of 0.873 and outperforms existing approaches by 20.6%. We further analyze factors affecting attack success, including query complexity, data type, data structure, and training settings. We also evaluate instruction-based, filter-based, and differential privacy-based defenses, which reduce performance but do not eliminate the risk. Our results highlight persistent privacy vulnerabilities in modern LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06423
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
Chen, Zeyuan
Ma, Yihan
Shen, Xinyue
Backes, Michael
Zhang, Yang
Cryptography and Security
Large language models (LLMs) show strong performance across many applications, but their ability to memorize and potentially reveal training data raises serious privacy concerns. We introduce the PopQuiz Attack, a black-box membership inference attack that tests whether a model can recall specific training examples. The core idea is to turn target data into quiz-style multiple-choice questions and infer membership from the model's answers. Across six widely used LLMs (GPT-3.5, GPT-4o, LLaMA2-7b, LLaMA2-13b, Mistral-7b, and Vicuna-7b) and four datasets, our method achieves an average ROC-AUC of 0.873 and outperforms existing approaches by 20.6%. We further analyze factors affecting attack success, including query complexity, data type, data structure, and training settings. We also evaluate instruction-based, filter-based, and differential privacy-based defenses, which reduce performance but do not eliminate the risk. Our results highlight persistent privacy vulnerabilities in modern LLMs.
title Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
topic Cryptography and Security
url https://arxiv.org/abs/2605.06423