Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915445668315136 |
|---|---|
| author | Jang, Yeonwoo Hossain, Shariqah Sreevatsa, Ashwin Cruz, Diogo |
| author_facet | Jang, Yeonwoo Hossain, Shariqah Sreevatsa, Ashwin Cruz, Diogo |
| contents | In this work, we demonstrate that certain machine unlearning methods may fail under straightforward prompt attacks. We systematically evaluate eight unlearning techniques across three model families using output-based, logit-based, and probe analysis to assess the extent to which supposedly unlearned knowledge can be retrieved. While methods like RMU and TAR exhibit robust unlearning, ELM remains vulnerable to specific prompt attacks (e.g., prepending Hindi filler text to the original prompt recovers 57.3% accuracy). Our logit analysis further indicates that unlearned models are unlikely to hide knowledge through changes in answer formatting, given the strong correlation between output and logit accuracy. These findings challenge prevailing assumptions about unlearning effectiveness and highlight the need for evaluation frameworks that can reliably distinguish between genuine knowledge removal and superficial output suppression. To facilitate further research, we publicly release our evaluation framework to easily evaluate prompting techniques to retrieve unlearned knowledge. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_10236 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Jang, Yeonwoo Hossain, Shariqah Sreevatsa, Ashwin Cruz, Diogo Cryptography and Security Artificial Intelligence Computation and Language Computers and Society Machine Learning In this work, we demonstrate that certain machine unlearning methods may fail under straightforward prompt attacks. We systematically evaluate eight unlearning techniques across three model families using output-based, logit-based, and probe analysis to assess the extent to which supposedly unlearned knowledge can be retrieved. While methods like RMU and TAR exhibit robust unlearning, ELM remains vulnerable to specific prompt attacks (e.g., prepending Hindi filler text to the original prompt recovers 57.3% accuracy). Our logit analysis further indicates that unlearned models are unlikely to hide knowledge through changes in answer formatting, given the strong correlation between output and logit accuracy. These findings challenge prevailing assumptions about unlearning effectiveness and highlight the need for evaluation frameworks that can reliably distinguish between genuine knowledge removal and superficial output suppression. To facilitate further research, we publicly release our evaluation framework to easily evaluate prompting techniques to retrieve unlearned knowledge. |
| title | Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods |
| topic | Cryptography and Security Artificial Intelligence Computation and Language Computers and Society Machine Learning |
| url | https://arxiv.org/abs/2506.10236 |