Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Shengyuan, Fu, Yiwei, Wu, Zhiwei Steven, Smith, Virginia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
di: Hu, Shengyuan, et al.
Pubblicazione: (2025)
di: Hu, Shengyuan, et al.
Pubblicazione: (2025)
Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026)
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026)
Layered Unlearning for Adversarial Relearning
di: Qian, Timothy, et al.
Pubblicazione: (2025)
di: Qian, Timothy, et al.
Pubblicazione: (2025)
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
di: Gao, Hongcheng, et al.
Pubblicazione: (2024)
di: Gao, Hongcheng, et al.
Pubblicazione: (2024)
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
di: Ha, SeungBum, et al.
Pubblicazione: (2025)
di: Ha, SeungBum, et al.
Pubblicazione: (2025)
Guardrail Baselines for Unlearning in LLMs
di: Thaker, Pratiksha, et al.
Pubblicazione: (2024)
di: Thaker, Pratiksha, et al.
Pubblicazione: (2024)
Efficient Unlearning through Maximizing Relearning Convergence Delay
di: Tran, Khoa, et al.
Pubblicazione: (2026)
di: Tran, Khoa, et al.
Pubblicazione: (2026)
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
di: Chen, Yiwei, et al.
Pubblicazione: (2025)
di: Chen, Yiwei, et al.
Pubblicazione: (2025)
COLUR: Confidence-Oriented Learning, Unlearning and Relearning with Noisy-Label Data for Model Restoration and Refinement
di: Sui, Zhihao, et al.
Pubblicazione: (2025)
di: Sui, Zhihao, et al.
Pubblicazione: (2025)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
di: Wu, Xiaoyu, et al.
Pubblicazione: (2025)
di: Wu, Xiaoyu, et al.
Pubblicazione: (2025)
FedCARE: Federated Unlearning with Conflict-Aware Projection and Relearning-Resistant Recovery
di: Li, Yue, et al.
Pubblicazione: (2026)
di: Li, Yue, et al.
Pubblicazione: (2026)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
di: Shin, Jeongjin, et al.
Pubblicazione: (2024)
di: Shin, Jeongjin, et al.
Pubblicazione: (2024)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
di: Sun, Guangzhi, et al.
Pubblicazione: (2025)
di: Sun, Guangzhi, et al.
Pubblicazione: (2025)
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
di: Fan, Chongyu, et al.
Pubblicazione: (2025)
di: Fan, Chongyu, et al.
Pubblicazione: (2025)
Exact Unlearning of Finetuning Data via Model Merging at Scale
di: Kuo, Kevin, et al.
Pubblicazione: (2025)
di: Kuo, Kevin, et al.
Pubblicazione: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
di: Xu, Xiaoyu, et al.
Pubblicazione: (2025)
di: Xu, Xiaoyu, et al.
Pubblicazione: (2025)
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
di: Thaker, Pratiksha, et al.
Pubblicazione: (2024)
di: Thaker, Pratiksha, et al.
Pubblicazione: (2024)
Enhancing One-run Privacy Auditing with Quantile Regression-Based Membership Inference
di: Liu, Terrance, et al.
Pubblicazione: (2025)
di: Liu, Terrance, et al.
Pubblicazione: (2025)
Reconstruction Attacks on Machine Unlearning: Simple Models are Vulnerable
di: Bertran, Martin, et al.
Pubblicazione: (2024)
di: Bertran, Martin, et al.
Pubblicazione: (2024)
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
di: Reisizadeh, Hadi, et al.
Pubblicazione: (2025)
di: Reisizadeh, Hadi, et al.
Pubblicazione: (2025)
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
di: Mahmud, Tamim Al, et al.
Pubblicazione: (2025)
di: Mahmud, Tamim Al, et al.
Pubblicazione: (2025)
FUNU: Boosting Machine Unlearning Efficiency by Filtering Unnecessary Unlearning
di: Li, Zitong, et al.
Pubblicazione: (2025)
di: Li, Zitong, et al.
Pubblicazione: (2025)
Unified Parameter-Efficient Unlearning for LLMs
di: Ding, Chenlu, et al.
Pubblicazione: (2024)
di: Ding, Chenlu, et al.
Pubblicazione: (2024)
Learning-Time Encoding Shapes Unlearning in LLMs
di: Wu, Ruihan, et al.
Pubblicazione: (2025)
di: Wu, Ruihan, et al.
Pubblicazione: (2025)
Leverage Unlearning to Sanitize LLMs
di: Boutet, Antoine, et al.
Pubblicazione: (2025)
di: Boutet, Antoine, et al.
Pubblicazione: (2025)
Online Learning and Unlearning
di: Hu, Yaxi, et al.
Pubblicazione: (2025)
di: Hu, Yaxi, et al.
Pubblicazione: (2025)
UCD: Unlearning in LLMs via Contrastive Decoding
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
di: Suriyakumar, Vinith M., et al.
Pubblicazione: (2025)
Targeted Unlearning with Single Layer Unlearning Gradient
di: Cai, Zikui, et al.
Pubblicazione: (2024)
di: Cai, Zikui, et al.
Pubblicazione: (2024)
Probing Knowledge Holes in Unlearned LLMs
di: Ko, Myeongseob, et al.
Pubblicazione: (2025)
di: Ko, Myeongseob, et al.
Pubblicazione: (2025)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
di: Kadhe, Swanand Ravindra, et al.
Pubblicazione: (2024)
di: Kadhe, Swanand Ravindra, et al.
Pubblicazione: (2024)
Unlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse Adapter
di: Zhong, Zhengyi, et al.
Pubblicazione: (2025)
di: Zhong, Zhengyi, et al.
Pubblicazione: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
di: Guo, Phillip, et al.
Pubblicazione: (2024)
di: Guo, Phillip, et al.
Pubblicazione: (2024)
Federated Graph Unlearning
di: Ai, Yuming, et al.
Pubblicazione: (2025)
di: Ai, Yuming, et al.
Pubblicazione: (2025)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
di: Qiu, Xinchi, et al.
Pubblicazione: (2024)
di: Qiu, Xinchi, et al.
Pubblicazione: (2024)
Conformal Unlearning: A New Paradigm for Unlearning in Conformal Predictors
di: Alkhatib, Yahya, et al.
Pubblicazione: (2025)
di: Alkhatib, Yahya, et al.
Pubblicazione: (2025)
An Illusion of Unlearning? Assessing Machine Unlearning Through Internal Representations
di: Gao, Yichen, et al.
Pubblicazione: (2026)
di: Gao, Yichen, et al.
Pubblicazione: (2026)
Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
di: Yang, Nakyeong, et al.
Pubblicazione: (2025)
di: Yang, Nakyeong, et al.
Pubblicazione: (2025)
Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
di: Cha, Sungmin, et al.
Pubblicazione: (2023)
di: Cha, Sungmin, et al.
Pubblicazione: (2023)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
di: Spohn, Philipp, et al.
Pubblicazione: (2025)
di: Spohn, Philipp, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
di: Hu, Shengyuan, et al.
Pubblicazione: (2025) -
Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures
di: Yoon, Sangyeon, et al.
Pubblicazione: (2026) -
Layered Unlearning for Adversarial Relearning
di: Qian, Timothy, et al.
Pubblicazione: (2025) -
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
di: Gao, Hongcheng, et al.
Pubblicazione: (2024) -
Unlearning's Blind Spots: Over-Unlearning and Prototypical Relearning Attack
di: Ha, SeungBum, et al.
Pubblicazione: (2025)