Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jang, Yeonwoo, Hossain, Shariqah, Sreevatsa, Ashwin, Cruz, Diogo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915445668315136
author Jang, Yeonwoo
Hossain, Shariqah
Sreevatsa, Ashwin
Cruz, Diogo
author_facet Jang, Yeonwoo
Hossain, Shariqah
Sreevatsa, Ashwin
Cruz, Diogo
contents In this work, we demonstrate that certain machine unlearning methods may fail under straightforward prompt attacks. We systematically evaluate eight unlearning techniques across three model families using output-based, logit-based, and probe analysis to assess the extent to which supposedly unlearned knowledge can be retrieved. While methods like RMU and TAR exhibit robust unlearning, ELM remains vulnerable to specific prompt attacks (e.g., prepending Hindi filler text to the original prompt recovers 57.3% accuracy). Our logit analysis further indicates that unlearned models are unlikely to hide knowledge through changes in answer formatting, given the strong correlation between output and logit accuracy. These findings challenge prevailing assumptions about unlearning effectiveness and highlight the need for evaluation frameworks that can reliably distinguish between genuine knowledge removal and superficial output suppression. To facilitate further research, we publicly release our evaluation framework to easily evaluate prompting techniques to retrieve unlearned knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10236
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
Jang, Yeonwoo
Hossain, Shariqah
Sreevatsa, Ashwin
Cruz, Diogo
Cryptography and Security
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
In this work, we demonstrate that certain machine unlearning methods may fail under straightforward prompt attacks. We systematically evaluate eight unlearning techniques across three model families using output-based, logit-based, and probe analysis to assess the extent to which supposedly unlearned knowledge can be retrieved. While methods like RMU and TAR exhibit robust unlearning, ELM remains vulnerable to specific prompt attacks (e.g., prepending Hindi filler text to the original prompt recovers 57.3% accuracy). Our logit analysis further indicates that unlearned models are unlikely to hide knowledge through changes in answer formatting, given the strong correlation between output and logit accuracy. These findings challenge prevailing assumptions about unlearning effectiveness and highlight the need for evaluation frameworks that can reliably distinguish between genuine knowledge removal and superficial output suppression. To facilitate further research, we publicly release our evaluation framework to easily evaluate prompting techniques to retrieve unlearned knowledge.
title Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2506.10236