Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fu, Haowei, Ni, Bo, Xu, Han, Liu, Kunpeng, Lin, Dan, Derr, Tyler
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918229567340544
author Fu, Haowei
Ni, Bo
Xu, Han
Liu, Kunpeng
Lin, Dan
Derr, Tyler
author_facet Fu, Haowei
Ni, Bo
Xu, Han
Liu, Kunpeng
Lin, Dan
Derr, Tyler
contents Retrieval-Augmented Generation (RAG) and Supervised Finetuning (SFT) have become the predominant paradigms for equipping Large Language Models (LLMs) with external knowledge for diverse, knowledge-intensive tasks. However, while such knowledge injection improves performance, it also exposes new attack surfaces. Membership Inference Attacks (MIAs), which aim to determine whether a given data sample was included in a model's training set, pose serious threats to privacy and trust in sensitive domains. To this end, we first systematically evaluate the vulnerability of RAG- and SFT-based LLMs to various MIAs. Then, to address the privacy risk, we further introduce a novel, model-agnostic defense framework, Ensemble Privacy Defense (EPD), which aggregates and evaluates the outputs of a knowledge-injected LLM, a base LLM, and a dedicated judge model to enhance resistance against MIAs. Comprehensive experiments show that, on average, EPD reduces MIA success by up to 27.8\% for SFT and 526.3\% for RAG compared to inference-time baseline, while maintaining answer quality.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03100
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
Fu, Haowei
Ni, Bo
Xu, Han
Liu, Kunpeng
Lin, Dan
Derr, Tyler
Cryptography and Security
Artificial Intelligence
Retrieval-Augmented Generation (RAG) and Supervised Finetuning (SFT) have become the predominant paradigms for equipping Large Language Models (LLMs) with external knowledge for diverse, knowledge-intensive tasks. However, while such knowledge injection improves performance, it also exposes new attack surfaces. Membership Inference Attacks (MIAs), which aim to determine whether a given data sample was included in a model's training set, pose serious threats to privacy and trust in sensitive domains. To this end, we first systematically evaluate the vulnerability of RAG- and SFT-based LLMs to various MIAs. Then, to address the privacy risk, we further introduce a novel, model-agnostic defense framework, Ensemble Privacy Defense (EPD), which aggregates and evaluates the outputs of a knowledge-injected LLM, a base LLM, and a dedicated judge model to enhance resistance against MIAs. Comprehensive experiments show that, on average, EPD reduces MIA success by up to 27.8\% for SFT and 526.3\% for RAG compared to inference-time baseline, while maintaining answer quality.
title Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2512.03100