Membership Inference Attacks for Retrieval Based In-Context Learning for Document Question Answering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kulkarni, Tejas, Koskela, Antti, Zumot, Laith
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913091864756224
author Kulkarni, Tejas
Koskela, Antti
Zumot, Laith
author_facet Kulkarni, Tejas
Koskela, Antti
Zumot, Laith
contents We show that remotely hosted applications employing in-context learning when augmented with a retrieval function to select in-context examples can be vulnerable to membership-inference attacks even when the service provider and users are separate parties. We propose two black-box membership inference attacks that exploit query text prefixes to distinguish member from non-member inputs. The first attack uses a reference model to estimate an otherwise unavailable loss metric. The second attack improves upon it by eliminating the reference model and instead computing a membership statistic through a simple but novel weighted-averaging scheme. Our comprehensive empirical evaluations consider a stricter case in which the adversary has a paraphrased version of the text in the queries and show that our attacks can exhibit stronger resilience to paraphrasing and outperform three prior attacks in many cases with small number of prefixes. We also adapt an existing ensemble prompting defense to our setting, demonstrating that it substantially mitigates the privacy leakage caused by our second attack.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04116
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Membership Inference Attacks for Retrieval Based In-Context Learning for Document Question Answering
Kulkarni, Tejas
Koskela, Antti
Zumot, Laith
Cryptography and Security
Machine Learning
We show that remotely hosted applications employing in-context learning when augmented with a retrieval function to select in-context examples can be vulnerable to membership-inference attacks even when the service provider and users are separate parties. We propose two black-box membership inference attacks that exploit query text prefixes to distinguish member from non-member inputs. The first attack uses a reference model to estimate an otherwise unavailable loss metric. The second attack improves upon it by eliminating the reference model and instead computing a membership statistic through a simple but novel weighted-averaging scheme. Our comprehensive empirical evaluations consider a stricter case in which the adversary has a paraphrased version of the text in the queries and show that our attacks can exhibit stronger resilience to paraphrasing and outperform three prior attacks in many cases with small number of prefixes. We also adapt an existing ensemble prompting defense to our setting, demonstrating that it substantially mitigates the privacy leakage caused by our second attack.
title Membership Inference Attacks for Retrieval Based In-Context Learning for Document Question Answering
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2605.04116