ROKA: Robust Knowledge Unlearning against Adversaries

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shin, Jinmyeong, Tapia, Joshua, Ferreira, Nicholas, Diaz, Gabriel, Daneshyari, Moayed, Jeon, Hyeran
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910035920027648
author Shin, Jinmyeong
Tapia, Joshua
Ferreira, Nicholas
Diaz, Gabriel
Daneshyari, Moayed
Jeon, Hyeran
author_facet Shin, Jinmyeong
Tapia, Joshua
Ferreira, Nicholas
Diaz, Gabriel
Daneshyari, Moayed
Jeon, Hyeran
contents The need for machine unlearning is critical for data privacy, yet existing methods often cause Knowledge Contamination by unintentionally damaging related knowledge. Such a degraded model performance after unlearning has been recently leveraged for new inference and backdoor attacks. Most studies design adversarial unlearning requests that require poisoning or duplicating training data. In this study, we introduce a new unlearning-induced attack model, namely indirect unlearning attack, which does not require data manipulation but exploits the consequence of knowledge contamination to perturb the model accuracy on security-critical predictions. To mitigate this attack, we introduce a theoretical framework that models neural networks as Neural Knowledge Systems. Based on this, we propose ROKA, a robust unlearning strategy centered on Neural Healing. Unlike conventional unlearning methods that only destroy information, ROKA constructively rebalances the model by nullifying the influence of forgotten data while strengthening its conceptual neighbors. To the best of our knowledge, our work is the first to provide a theoretical guarantee for knowledge preservation during unlearning. Evaluations on various large models, including vision transformers, multi-modal models, and large language models, show that ROKA effectively unlearns targets while preserving, or even enhancing, the accuracy of retained data, thereby mitigating the indirect unlearning attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_00436
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ROKA: Robust Knowledge Unlearning against Adversaries
Shin, Jinmyeong
Tapia, Joshua
Ferreira, Nicholas
Diaz, Gabriel
Daneshyari, Moayed
Jeon, Hyeran
Machine Learning
Artificial Intelligence
The need for machine unlearning is critical for data privacy, yet existing methods often cause Knowledge Contamination by unintentionally damaging related knowledge. Such a degraded model performance after unlearning has been recently leveraged for new inference and backdoor attacks. Most studies design adversarial unlearning requests that require poisoning or duplicating training data. In this study, we introduce a new unlearning-induced attack model, namely indirect unlearning attack, which does not require data manipulation but exploits the consequence of knowledge contamination to perturb the model accuracy on security-critical predictions. To mitigate this attack, we introduce a theoretical framework that models neural networks as Neural Knowledge Systems. Based on this, we propose ROKA, a robust unlearning strategy centered on Neural Healing. Unlike conventional unlearning methods that only destroy information, ROKA constructively rebalances the model by nullifying the influence of forgotten data while strengthening its conceptual neighbors. To the best of our knowledge, our work is the first to provide a theoretical guarantee for knowledge preservation during unlearning. Evaluations on various large models, including vision transformers, multi-modal models, and large language models, show that ROKA effectively unlearns targets while preserving, or even enhancing, the accuracy of retained data, thereby mitigating the indirect unlearning attacks.
title ROKA: Robust Knowledge Unlearning against Adversaries
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.00436