Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kassem, Aly M., Mahmoud, Omar, Mireshghallah, Niloofar, Kim, Hyunwoo, Tsvetkov, Yulia, Choi, Yejin, Saad, Sherif, Rana, Santu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910819330031616
author Kassem, Aly M.
Mahmoud, Omar
Mireshghallah, Niloofar
Kim, Hyunwoo
Tsvetkov, Yulia
Choi, Yejin
Saad, Sherif
Rana, Santu
author_facet Kassem, Aly M.
Mahmoud, Omar
Mireshghallah, Niloofar
Kim, Hyunwoo
Tsvetkov, Yulia
Choi, Yejin
Saad, Sherif
Rana, Santu
contents In this paper, we introduce a black-box prompt optimization method that uses an attacker LLM agent to uncover higher levels of memorization in a victim agent, compared to what is revealed by prompting the target model with the training data directly, which is the dominant approach of quantifying memorization in LLMs. We use an iterative rejection-sampling optimization process to find instruction-based prompts with two main characteristics: (1) minimal overlap with the training data to avoid presenting the solution directly to the model, and (2) maximal overlap between the victim model's output and the training data, aiming to induce the victim to spit out training data. We observe that our instruction-based prompts generate outputs with 23.7% higher overlap with training data compared to the baseline prefix-suffix measurements. Our findings show that (1) instruction-tuned models can expose pre-training data as much as their base-models, if not more so, (2) contexts other than the original training data can lead to leakage, and (3) using instructions proposed by other LLMs can open a new avenue of automated attacks that we should further study and explore. The code can be found at https://github.com/Alymostafa/Instruction_based_attack .
format Preprint
id arxiv_https___arxiv_org_abs_2403_04801
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
Kassem, Aly M.
Mahmoud, Omar
Mireshghallah, Niloofar
Kim, Hyunwoo
Tsvetkov, Yulia
Choi, Yejin
Saad, Sherif
Rana, Santu
Computation and Language
In this paper, we introduce a black-box prompt optimization method that uses an attacker LLM agent to uncover higher levels of memorization in a victim agent, compared to what is revealed by prompting the target model with the training data directly, which is the dominant approach of quantifying memorization in LLMs. We use an iterative rejection-sampling optimization process to find instruction-based prompts with two main characteristics: (1) minimal overlap with the training data to avoid presenting the solution directly to the model, and (2) maximal overlap between the victim model's output and the training data, aiming to induce the victim to spit out training data. We observe that our instruction-based prompts generate outputs with 23.7% higher overlap with training data compared to the baseline prefix-suffix measurements. Our findings show that (1) instruction-tuned models can expose pre-training data as much as their base-models, if not more so, (2) contexts other than the original training data can lead to leakage, and (3) using instructions proposed by other LLMs can open a new avenue of automated attacks that we should further study and explore. The code can be found at https://github.com/Alymostafa/Instruction_based_attack .
title Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
topic Computation and Language
url https://arxiv.org/abs/2403.04801