Improving LLM Unlearning Robustness via Random Perturbations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huu-Tien, Dang, Thanh-Tung, Hoang, Bui, Anh, Nguyen, Minh-Phuong, Nguyen, Le-Minh, Inoue, Naoya
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913043290521600
author Huu-Tien, Dang
Thanh-Tung, Hoang
Bui, Anh
Nguyen, Minh-Phuong
Nguyen, Le-Minh
Inoue, Naoya
author_facet Huu-Tien, Dang
Thanh-Tung, Hoang
Bui, Anh
Nguyen, Minh-Phuong
Nguyen, Le-Minh
Inoue, Naoya
contents Here, we show that current LLM unlearning methods inherently reduce models' robustness, causing them to misbehave even when a single non-adversarial forget-token is present in the retain-query. Toward understanding underlying causes, we propose a novel theoretical framework that reframes the unlearning process as a backdoor attack and defense problem: we formulate how the forgetting process inadvertently learns to align forget-tokens (backdoor triggers) with the target-representations (target labels). As a result, forget-tokens act as backdoor triggers that, when activated in retain-queries, cause disruptions in unlearned models' behaviors, similar to successful backdoor attacks. The sense that, LLM unlearning methods themselves poison the model, make it more vulnerable to forget-tokens, and hide rather than erase target knowledge, describes their true mechanism. To mitigate the vulnerability caused by the forgetting process, we reinterpret the retaining process as a backdoor defense and propose Random Noise Augmentation (RNA), a lightweight, model and method-agnostic approach with theoretical guarantees for improving the robustness of unlearned models. Extensive experiments demonstrate that RNA significantly improves the robustness of unlearned models while preserving forget and retain performances. This backdoor attack-defense framework offers insights into the mechanism of unlearning that can shed light on future research directions for improving unlearning robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2501_19202
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving LLM Unlearning Robustness via Random Perturbations
Huu-Tien, Dang
Thanh-Tung, Hoang
Bui, Anh
Nguyen, Minh-Phuong
Nguyen, Le-Minh
Inoue, Naoya
Computation and Language
Here, we show that current LLM unlearning methods inherently reduce models' robustness, causing them to misbehave even when a single non-adversarial forget-token is present in the retain-query. Toward understanding underlying causes, we propose a novel theoretical framework that reframes the unlearning process as a backdoor attack and defense problem: we formulate how the forgetting process inadvertently learns to align forget-tokens (backdoor triggers) with the target-representations (target labels). As a result, forget-tokens act as backdoor triggers that, when activated in retain-queries, cause disruptions in unlearned models' behaviors, similar to successful backdoor attacks. The sense that, LLM unlearning methods themselves poison the model, make it more vulnerable to forget-tokens, and hide rather than erase target knowledge, describes their true mechanism. To mitigate the vulnerability caused by the forgetting process, we reinterpret the retaining process as a backdoor defense and propose Random Noise Augmentation (RNA), a lightweight, model and method-agnostic approach with theoretical guarantees for improving the robustness of unlearned models. Extensive experiments demonstrate that RNA significantly improves the robustness of unlearned models while preserving forget and retain performances. This backdoor attack-defense framework offers insights into the mechanism of unlearning that can shed light on future research directions for improving unlearning robustness.
title Improving LLM Unlearning Robustness via Random Perturbations
topic Computation and Language
url https://arxiv.org/abs/2501.19202