From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
Fuente:
arXiv
Saved in:
| Main Authors: | Siddiqui, Shoaib Ahmed, Weller, Adrian, Krueger, David, Dziugaite, Gintare Karolina, Mozer, Michael Curtis, Triantafillou, Eleni |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Capability Control Should be a Separate Goal From Alignment
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026)
Data Selection for Transfer Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
Improved Localized Machine Unlearning Through the Lens of Memorization
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
Leveraging Per-Instance Privacy for Machine Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)
by: Guo, Phillip, et al.
Published: (2024)
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2026)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2026)
Unlearning in- vs. out-of-distribution data in LLMs under gradient-based method
by: Baluta, Teodora, et al.
Published: (2024)
by: Baluta, Teodora, et al.
Published: (2024)
The Topological Trouble With Transformers
by: Mozer, Michael C., et al.
Published: (2026)
by: Mozer, Michael C., et al.
Published: (2026)
The Non-Local Model Merging Problem: Permutation Symmetries and Variance Collapse
by: Sharma, Ekansh, et al.
Published: (2024)
by: Sharma, Ekansh, et al.
Published: (2024)
Leveraging Function Space Aggregation for Federated Learning at Scale
by: Dhawan, Nikita, et al.
Published: (2023)
by: Dhawan, Nikita, et al.
Published: (2023)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
by: Yang, Yu, et al.
Published: (2023)
by: Yang, Yu, et al.
Published: (2023)
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
by: Inane, Ahmed Mehdi, et al.
Published: (2026)
by: Inane, Ahmed Mehdi, et al.
Published: (2026)
Dataset Difficulty and the Role of Inductive Bias
by: Kwok, Devin, et al.
Published: (2024)
by: Kwok, Devin, et al.
Published: (2024)
Protecting against simultaneous data poisoning attacks
by: Alex, Neel, et al.
Published: (2024)
by: Alex, Neel, et al.
Published: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
Soup to go: mitigating forgetting during continual learning with model averaging
by: Kleiman, Anat, et al.
Published: (2025)
by: Kleiman, Anat, et al.
Published: (2025)
Redirection for Erasing Memory (REM): Towards a universal unlearning method for corrupted data
by: Schoepf, Stefan, et al.
Published: (2025)
by: Schoepf, Stefan, et al.
Published: (2025)
Information Complexity of Stochastic Convex Optimization: Applications to Generalization and Memorization
by: Attias, Idan, et al.
Published: (2024)
by: Attias, Idan, et al.
Published: (2024)
Blockwise Self-Supervised Learning at Scale
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2023)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Is your algorithm unlearning or untraining?
by: Triantafillou, Eleni, et al.
Published: (2026)
by: Triantafillou, Eleni, et al.
Published: (2026)
Simultaneous linear connectivity of neural networks modulo permutation
by: Sharma, Ekansh, et al.
Published: (2024)
by: Sharma, Ekansh, et al.
Published: (2024)
Are we making progress in unlearning? Findings from the first NeurIPS unlearning competition
by: Triantafillou, Eleni, et al.
Published: (2024)
by: Triantafillou, Eleni, et al.
Published: (2024)
On Evaluating LLMs' Capabilities as Functional Approximators: A Bayesian Perspective
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
SSFL: Discovering Sparse Unified Subnetworks at Initialization for Efficient Federated Learning
by: Ohib, Riyasat, et al.
Published: (2024)
by: Ohib, Riyasat, et al.
Published: (2024)
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
by: Jin, Tian, et al.
Published: (2025)
by: Jin, Tian, et al.
Published: (2025)
On Traceability in $\ell_p$ Stochastic Convex Optimization
by: Voitovych, Sasha, et al.
Published: (2025)
by: Voitovych, Sasha, et al.
Published: (2025)
Continual Learning in Vision-Language Models via Aligned Model Merging
by: Sokar, Ghada, et al.
Published: (2025)
by: Sokar, Ghada, et al.
Published: (2025)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery
by: Gopalakrishnan, Anand, et al.
Published: (2024)
by: Gopalakrishnan, Anand, et al.
Published: (2024)
Benchmarking Unlearning for Vision Transformers
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Improving Discrete Optimisation Via Decoupled Straight-Through Estimator
by: Shah, Rushi, et al.
Published: (2024)
by: Shah, Rushi, et al.
Published: (2024)
Torque-Aware Momentum
by: Malviya, Pranshu, et al.
Published: (2024)
by: Malviya, Pranshu, et al.
Published: (2024)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
What makes unlearning hard and what to do about it
by: Zhao, Kairan, et al.
Published: (2024)
by: Zhao, Kairan, et al.
Published: (2024)
A deeper look at depth pruning of LLMs
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
Exploring the design space of deep-learning-based weather forecasting systems
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2024)
To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language Models
by: Barbulescu, George-Octavian, et al.
Published: (2024)
by: Barbulescu, George-Octavian, et al.
Published: (2024)
Similar Items
-
Position: Capability Control Should be a Separate Goal From Alignment
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026) -
Data Selection for Transfer Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024) -
Improved Localized Machine Unlearning Through the Lens of Memorization
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024) -
Leveraging Per-Instance Privacy for Machine Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025) -
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)