MRMMIA: Membership Inference Attacks on Memory in Chat Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Kai, Pang, Yan, Wang, Tianhao
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911723445813248
author Chen, Kai
Pang, Yan
Wang, Tianhao
author_facet Chen, Kai
Pang, Yan
Wang, Tianhao
contents Membership inference attacks (MIAs) test whether a target data record belongs to a system's private data, and have become a standard tool to measure privacy leakage in machine learning systems. Prior work has primarily focused on training corpora or retrieval databases. However, MIAs against agent memory have received less attention, even though such memory can contain sensitive user-agent interactions, retrieved facts, and user preferences. Therefore, in this work, we focus on chat agent memory MIAs, where an adversary infers whether a candidate memory unit belongs to the chat agent's memory store. We propose Multi-Recall Memory MIA (MRMMIA), a unified attack that utilizes multiple recall probes to the agent to extract the membership signal across black-box, gray-box, and white-box settings. Our experiments demonstrate that MRMMIA consistently outperforms baselines. Our results expose the privacy risk in agents and provide an initial evaluation framework for membership leakage in chat-agent memory systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27825
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MRMMIA: Membership Inference Attacks on Memory in Chat Agents
Chen, Kai
Pang, Yan
Wang, Tianhao
Cryptography and Security
Machine Learning
Membership inference attacks (MIAs) test whether a target data record belongs to a system's private data, and have become a standard tool to measure privacy leakage in machine learning systems. Prior work has primarily focused on training corpora or retrieval databases. However, MIAs against agent memory have received less attention, even though such memory can contain sensitive user-agent interactions, retrieved facts, and user preferences. Therefore, in this work, we focus on chat agent memory MIAs, where an adversary infers whether a candidate memory unit belongs to the chat agent's memory store. We propose Multi-Recall Memory MIA (MRMMIA), a unified attack that utilizes multiple recall probes to the agent to extract the membership signal across black-box, gray-box, and white-box settings. Our experiments demonstrate that MRMMIA consistently outperforms baselines. Our results expose the privacy risk in agents and provide an initial evaluation framework for membership leakage in chat-agent memory systems.
title MRMMIA: Membership Inference Attacks on Memory in Chat Agents
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2605.27825