Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Cheng, Wang, Yiwei, Hooi, Bryan, Cai, Yujun, Peng, Nanyun, Chang, Kai-Wei
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915103664766976
author Wang, Cheng
Wang, Yiwei
Hooi, Bryan
Cai, Yujun
Peng, Nanyun
Chang, Kai-Wei
author_facet Wang, Cheng
Wang, Yiwei
Hooi, Bryan
Cai, Yujun
Peng, Nanyun
Chang, Kai-Wei
contents The training data in large language models is key to their success, but it also presents privacy and security risks, as it may contain sensitive information. Detecting pre-training data is crucial for mitigating these concerns. Existing methods typically analyze target text in isolation or solely with non-member contexts, overlooking potential insights from simultaneously considering both member and non-member contexts. While previous work suggested that member contexts provide little information due to the minor distributional shift they induce, our analysis reveals that these subtle shifts can be effectively leveraged when contrasted with non-member contexts. In this paper, we propose Con-ReCall, a novel approach that leverages the asymmetric distributional shifts induced by member and non-member contexts through contrastive decoding, amplifying subtle differences to enhance membership inference. Extensive empirical evaluations demonstrate that Con-ReCall achieves state-of-the-art performance on the WikiMIA benchmark and is robust against various text manipulation techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2409_03363
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
Wang, Cheng
Wang, Yiwei
Hooi, Bryan
Cai, Yujun
Peng, Nanyun
Chang, Kai-Wei
Computation and Language
The training data in large language models is key to their success, but it also presents privacy and security risks, as it may contain sensitive information. Detecting pre-training data is crucial for mitigating these concerns. Existing methods typically analyze target text in isolation or solely with non-member contexts, overlooking potential insights from simultaneously considering both member and non-member contexts. While previous work suggested that member contexts provide little information due to the minor distributional shift they induce, our analysis reveals that these subtle shifts can be effectively leveraged when contrasted with non-member contexts. In this paper, we propose Con-ReCall, a novel approach that leverages the asymmetric distributional shifts induced by member and non-member contexts through contrastive decoding, amplifying subtle differences to enhance membership inference. Extensive empirical evaluations demonstrate that Con-ReCall achieves state-of-the-art performance on the WikiMIA benchmark and is robust against various text manipulation techniques.
title Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
topic Computation and Language
url https://arxiv.org/abs/2409.03363