Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing
Fuente:
arXiv
Saved in:
| Main Authors: | Kassem, Aly M., Shi, Zhuan, Rostamzadeh, Negar, Farnadi, Golnoosh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
by: Kassem, Aly, et al.
Published: (2026)
by: Kassem, Aly, et al.
Published: (2026)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026)
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
by: Farnadi, Golnoosh, et al.
Published: (2024)
by: Farnadi, Golnoosh, et al.
Published: (2024)
LoRA Provides Differential Privacy by Design via Random Sketching
by: Malekmohammadi, Saber, et al.
Published: (2024)
by: Malekmohammadi, Saber, et al.
Published: (2024)
Understanding Intrinsic Socioeconomic Biases in Large Language Models
by: Arzaghi, Mina, et al.
Published: (2024)
by: Arzaghi, Mina, et al.
Published: (2024)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
by: More, Yash, et al.
Published: (2024)
by: More, Yash, et al.
Published: (2024)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
by: Arzaghi', 'Mina, et al.
Published: (2025)
by: Arzaghi', 'Mina, et al.
Published: (2025)
What Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models
by: Humayun, Ahmed Imtiaz, et al.
Published: (2024)
by: Humayun, Ahmed Imtiaz, et al.
Published: (2024)
LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
by: Liu, Yezi, et al.
Published: (2025)
by: Liu, Yezi, et al.
Published: (2025)
Data as a Lever: A Neighbouring Datasets Perspective on Predictive Multiplicity
by: Ganesh, Prakhar, et al.
Published: (2025)
by: Ganesh, Prakhar, et al.
Published: (2025)
Dissecting Fine-Tuning Unlearning in Large Language Models
by: Hong, Yihuai, et al.
Published: (2024)
by: Hong, Yihuai, et al.
Published: (2024)
Wasserstein Distributionally Robust Optimization Through the Lens of Structural Causal Models and Individual Fairness
by: Ehyaei, Ahmad-Reza, et al.
Published: (2025)
by: Ehyaei, Ahmad-Reza, et al.
Published: (2025)
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Systemizing Multiplicity: The Curious Case of Arbitrariness in Machine Learning
by: Ganesh, Prakhar, et al.
Published: (2025)
by: Ganesh, Prakhar, et al.
Published: (2025)
Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity
by: Ganesh, Prakhar, et al.
Published: (2026)
by: Ganesh, Prakhar, et al.
Published: (2026)
What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning
by: Zhu, Yuchang, et al.
Published: (2025)
by: Zhu, Yuchang, et al.
Published: (2025)
Scaling Sparse Fine-Tuning to Large Language Models
by: Ansell, Alan, et al.
Published: (2024)
by: Ansell, Alan, et al.
Published: (2024)
From Fragile to Certified: Wasserstein Audits of Group Fairness Under Distribution Shift
by: Ehyaei, Ahmad-Reza, et al.
Published: (2025)
by: Ehyaei, Ahmad-Reza, et al.
Published: (2025)
Designing Ambiguity Sets for Distributionally Robust Optimization Using Structural Causal Optimal Transport
by: Ehyaei, Ahmad-Reza, et al.
Published: (2025)
by: Ehyaei, Ahmad-Reza, et al.
Published: (2025)
Understanding the role of depth in the neural tangent kernel for overparameterized neural networks
by: St-Arnaud, William, et al.
Published: (2025)
by: St-Arnaud, William, et al.
Published: (2025)
A General Framework to Enhance Fine-tuning-based LLM Unlearning
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
by: Zheng, Estelle, et al.
Published: (2025)
by: Zheng, Estelle, et al.
Published: (2025)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
by: Guo, Song, et al.
Published: (2024)
by: Guo, Song, et al.
Published: (2024)
Fairness in Federated Learning: Fairness for Whom?
by: Taik, Afaf, et al.
Published: (2025)
by: Taik, Afaf, et al.
Published: (2025)
Advancing Cultural Inclusivity: Optimizing Embedding Spaces for Balanced Music Recommendations
by: Moradi, Armin, et al.
Published: (2024)
by: Moradi, Armin, et al.
Published: (2024)
Balancing Act: Constraining Disparate Impact in Sparse Models
by: Hashemizadeh, Meraj, et al.
Published: (2023)
by: Hashemizadeh, Meraj, et al.
Published: (2023)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
by: Liu, James, et al.
Published: (2024)
by: Liu, James, et al.
Published: (2024)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
by: Spohn, Philipp, et al.
Published: (2025)
by: Spohn, Philipp, et al.
Published: (2025)
Causal Fair Metric: Bridging Causality, Individual Fairness, and Adversarial Robustness
by: Ehyaei, Ahmad-Reza, et al.
Published: (2023)
by: Ehyaei, Ahmad-Reza, et al.
Published: (2023)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
LLM Unlearning with LLM Beliefs
by: Li, Kemou, et al.
Published: (2025)
by: Li, Kemou, et al.
Published: (2025)
Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
by: Liu, Yong, et al.
Published: (2024)
by: Liu, Yong, et al.
Published: (2024)
Differentially Private Clustered Federated Learning
by: Malekmohammadi, Saber, et al.
Published: (2024)
by: Malekmohammadi, Saber, et al.
Published: (2024)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
by: Yamashita, Tomoya, et al.
Published: (2025)
by: Yamashita, Tomoya, et al.
Published: (2025)
Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning
by: Abdalla, M. H. I., et al.
Published: (2025)
by: Abdalla, M. H. I., et al.
Published: (2025)
Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities
by: Dang, Tien, et al.
Published: (2026)
by: Dang, Tien, et al.
Published: (2026)
Multilingual Hallucination Gaps in Large Language Models
by: Chataigner, Cléa, et al.
Published: (2024)
by: Chataigner, Cléa, et al.
Published: (2024)
Simple LLM Baselines are Competitive for Model Diffing
by: Kempf, Elias, et al.
Published: (2026)
by: Kempf, Elias, et al.
Published: (2026)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning
by: Li, Cen-Jhih, et al.
Published: (2025)
by: Li, Cen-Jhih, et al.
Published: (2025)
Similar Items
-
Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes
by: Kassem, Aly, et al.
Published: (2026) -
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
by: Farashah, Alireza Dehghanpour, et al.
Published: (2026) -
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
by: Farnadi, Golnoosh, et al.
Published: (2024) -
LoRA Provides Differential Privacy by Design via Random Sketching
by: Malekmohammadi, Saber, et al.
Published: (2024) -
Understanding Intrinsic Socioeconomic Biases in Large Language Models
by: Arzaghi, Mina, et al.
Published: (2024)