Membership Inference Attack against Large Language Model-based Recommendation Systems: A New Distillation-based Paradigm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cuihong, Li, Xiaowen, Huang, Chuanhuan, Yin, Jitao, Sang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915873259782144
author Cuihong, Li
Xiaowen, Huang
Chuanhuan, Yin
Jitao, Sang
author_facet Cuihong, Li
Xiaowen, Huang
Chuanhuan, Yin
Jitao, Sang
contents Membership Inference Attack (MIA) aims to determine whether a specific data sample was included in the training dataset of a target model. Traditional MIA approaches rely on shadow models to mimic target model behavior, but their effectiveness diminishes for Large Language Model (LLM)-based recommendation systems due to the scale and complexity of training data. This paper introduces a novel knowledge distillation-based MIA paradigm tailored for LLM-based recommendation systems. Our method constructs a reference model via distillation, applying distinct strategies for member and non-member data to enhance discriminative capabilities. The paradigm extracts fused features (e.g., confidence, entropy, loss, and hidden layer vectors) from the reference model to train an attack model, overcoming limitations of individual features. Extensive experiments on extended datasets (Last.FM, MovieLens, Book-Crossing, Delicious) and diverse LLMs (T5, GPT-2, LLaMA3) demonstrate that our approach significantly outperforms shadow model-based MIAs and individual-feature baselines. The results show its practicality for privacy attacks in LLM-driven recommender systems.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14763
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Membership Inference Attack against Large Language Model-based Recommendation Systems: A New Distillation-based Paradigm
Cuihong, Li
Xiaowen, Huang
Chuanhuan, Yin
Jitao, Sang
Information Retrieval
Artificial Intelligence
Membership Inference Attack (MIA) aims to determine whether a specific data sample was included in the training dataset of a target model. Traditional MIA approaches rely on shadow models to mimic target model behavior, but their effectiveness diminishes for Large Language Model (LLM)-based recommendation systems due to the scale and complexity of training data. This paper introduces a novel knowledge distillation-based MIA paradigm tailored for LLM-based recommendation systems. Our method constructs a reference model via distillation, applying distinct strategies for member and non-member data to enhance discriminative capabilities. The paradigm extracts fused features (e.g., confidence, entropy, loss, and hidden layer vectors) from the reference model to train an attack model, overcoming limitations of individual features. Extensive experiments on extended datasets (Last.FM, MovieLens, Book-Crossing, Delicious) and diverse LLMs (T5, GPT-2, LLaMA3) demonstrate that our approach significantly outperforms shadow model-based MIAs and individual-feature baselines. The results show its practicality for privacy attacks in LLM-driven recommender systems.
title Membership Inference Attack against Large Language Model-based Recommendation Systems: A New Distillation-based Paradigm
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2511.14763