Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xiaoyu, Yue, Xiang, Liu, Yang, Ye, Qingqing, Zheng, Huadi, Hu, Peizhao, Du, Minxin, Hu, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
FIT to Forget: Robust Continual Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
Machine Unlearning of Pre-trained Large Language Models
by: Yao, Jin, et al.
Published: (2024)
by: Yao, Jin, et al.
Published: (2024)
MER-Inspector: Assessing model extraction risks from an attack-agnostic perspective
by: Zhang, Xinwei, et al.
Published: (2025)
by: Zhang, Xinwei, et al.
Published: (2025)
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
by: Hu, Hongsheng, et al.
Published: (2024)
by: Hu, Hongsheng, et al.
Published: (2024)
Textual Unlearning Gives a False Sense of Unlearning
by: Du, Jiacheng, et al.
Published: (2024)
by: Du, Jiacheng, et al.
Published: (2024)
A Sample-Level Evaluation and Generative Framework for Model Inversion Attacks
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Privacy for Free: Leveraging Local Differential Privacy Perturbed Data from Multiple Services
by: Du, Rong, et al.
Published: (2025)
by: Du, Rong, et al.
Published: (2025)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
by: Wu, Xiaoyu, et al.
Published: (2025)
by: Wu, Xiaoyu, et al.
Published: (2025)
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge?
by: Wang, Shang, et al.
Published: (2024)
by: Wang, Shang, et al.
Published: (2024)
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
by: Liang, Zi, et al.
Published: (2025)
by: Liang, Zi, et al.
Published: (2025)
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
by: Gao, Hongcheng, et al.
Published: (2024)
by: Gao, Hongcheng, et al.
Published: (2024)
On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning
by: Ye, Xiaotian, et al.
Published: (2026)
by: Ye, Xiaotian, et al.
Published: (2026)
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models
by: Liang, Zi, et al.
Published: (2024)
by: Liang, Zi, et al.
Published: (2024)
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
by: Liang, Zi, et al.
Published: (2025)
by: Liang, Zi, et al.
Published: (2025)
UCD: Unlearning in LLMs via Contrastive Decoding
by: Suriyakumar, Vinith M., et al.
Published: (2025)
by: Suriyakumar, Vinith M., et al.
Published: (2025)
Interactive Trimming against Evasive Online Data Manipulation Attacks: A Game-Theoretic Approach
by: Fu, Yue, et al.
Published: (2024)
by: Fu, Yue, et al.
Published: (2024)
LLM Unlearning Should Be Form-Independent
by: Ye, Xiaotian, et al.
Published: (2025)
by: Ye, Xiaotian, et al.
Published: (2025)
Really Unlearned? Verifying Machine Unlearning via Influential Sample Pairs
by: Xu, Heng, et al.
Published: (2024)
by: Xu, Heng, et al.
Published: (2024)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Contrastive Unlearning: A Contrastive Approach to Machine Unlearning
by: Lee, Hong kyu, et al.
Published: (2024)
by: Lee, Hong kyu, et al.
Published: (2024)
Unlearn and Burn: Adversarial Machine Unlearning Requests Destroy Model Accuracy
by: Huang, Yangsibo, et al.
Published: (2024)
by: Huang, Yangsibo, et al.
Published: (2024)
Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection
by: Liang, Zi, et al.
Published: (2026)
by: Liang, Zi, et al.
Published: (2026)
"Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation
by: Liang, Zi, et al.
Published: (2024)
by: Liang, Zi, et al.
Published: (2024)
Adversarial Machine Unlearning
by: Di, Zonglin, et al.
Published: (2024)
by: Di, Zonglin, et al.
Published: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024)
by: Łucki, Jakub, et al.
Published: (2024)
LLMs Can Unlearn Refusal with Only 1,000 Benign Samples
by: Guo, Yangyang, et al.
Published: (2026)
by: Guo, Yangyang, et al.
Published: (2026)
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
by: Liu, Ziyao, et al.
Published: (2024)
by: Liu, Ziyao, et al.
Published: (2024)
Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs
by: P, Dinesh Srivasthav, et al.
Published: (2026)
by: P, Dinesh Srivasthav, et al.
Published: (2026)
Federated Heavy Hitter Analytics with Local Differential Privacy
by: Zhang, Yuemin, et al.
Published: (2024)
by: Zhang, Yuemin, et al.
Published: (2024)
Machine Unlearning with Minimal Gradient Dependence for High Unlearning Ratios
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
When Kids Mode Isn't For Kids: Investigating TikTok's "Under 13 Experience"
by: Figueira, Olivia, et al.
Published: (2025)
by: Figueira, Olivia, et al.
Published: (2025)
Machine Unlearning in Large Language Models
by: Chen, Kongyang, et al.
Published: (2024)
by: Chen, Kongyang, et al.
Published: (2024)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
Split Unlearning
by: Yu, Guangsheng, et al.
Published: (2023)
by: Yu, Guangsheng, et al.
Published: (2023)
Similar Items
-
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2025) -
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
by: Xu, Xiaoyu, et al.
Published: (2026) -
FIT to Forget: Robust Continual Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2026) -
When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents
by: Xu, Xiaoyu, et al.
Published: (2026) -
Machine Unlearning of Pre-trained Large Language Models
by: Yao, Jin, et al.
Published: (2024)