Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sondej, Filip, Yang, Yushi, Kniejski, Mikołaj, Windys, Marcel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
Explainable LLM Unlearning Through Reasoning
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
LLM Unlearning Without an Expert Curated Dataset
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Geometric-disentangelment Unlearning
von: Zhou, Duo, et al.
Veröffentlicht: (2025)
von: Zhou, Duo, et al.
Veröffentlicht: (2025)
Large Language Model Unlearning
von: Yao, Yuanshun, et al.
Veröffentlicht: (2023)
von: Yao, Yuanshun, et al.
Veröffentlicht: (2023)
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
Measuring the Depth of LLM Unlearning via Activation Patching
von: Lee, Jaeung, et al.
Veröffentlicht: (2026)
von: Lee, Jaeung, et al.
Veröffentlicht: (2026)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024)
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
von: Xie, Linxi, et al.
Veröffentlicht: (2025)
von: Xie, Linxi, et al.
Veröffentlicht: (2025)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
von: Qiu, Xinchi, et al.
Veröffentlicht: (2024)
von: Qiu, Xinchi, et al.
Veröffentlicht: (2024)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Hierarchical Federated Unlearning for Large Language Models
von: Zhong, Yisheng, et al.
Veröffentlicht: (2025)
von: Zhong, Yisheng, et al.
Veröffentlicht: (2025)
CURE:Circuit-Aware Unlearning for LLM-based Recommendation
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024)
von: Veldanda, Akshaj Kumar, et al.
Veröffentlicht: (2024)
Textual Unlearning Gives a False Sense of Unlearning
von: Du, Jiacheng, et al.
Veröffentlicht: (2024)
von: Du, Jiacheng, et al.
Veröffentlicht: (2024)
CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
von: Kocak, Aysenur, et al.
Veröffentlicht: (2025)
von: Kocak, Aysenur, et al.
Veröffentlicht: (2025)
A Neuro-inspired Interpretation of Unlearning in Large Language Models through Sample-level Unlearning Difficulty
von: Feng, Xiaohua, et al.
Veröffentlicht: (2025)
von: Feng, Xiaohua, et al.
Veröffentlicht: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
Tool Unlearning for Tool-Augmented LLMs
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
EvoMU: Evolutionary Machine Unlearning
von: Batorski, Pawel, et al.
Veröffentlicht: (2026)
von: Batorski, Pawel, et al.
Veröffentlicht: (2026)
Offset Unlearning for Large Language Models
von: Huang, James Y., et al.
Veröffentlicht: (2024)
von: Huang, James Y., et al.
Veröffentlicht: (2024)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
von: Shumailov, Ilia, et al.
Veröffentlicht: (2024)
von: Shumailov, Ilia, et al.
Veröffentlicht: (2024)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2025)
Large Language Model Unlearning via Embedding-Corrupted Prompts
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
AdvAnchor: Enhancing Diffusion Model Unlearning with Adversarial Anchors
von: Zhao, Mengnan, et al.
Veröffentlicht: (2024)
von: Zhao, Mengnan, et al.
Veröffentlicht: (2024)
Does Machine Unlearning Truly Remove Knowledge?
von: Chen, Haokun, et al.
Veröffentlicht: (2025)
von: Chen, Haokun, et al.
Veröffentlicht: (2025)
Soft Prompting for Unlearning in Large Language Models
von: Bhaila, Karuna, et al.
Veröffentlicht: (2024)
von: Bhaila, Karuna, et al.
Veröffentlicht: (2024)
Machine Unlearning in Generative AI: A Survey
von: Liu, Zheyuan, et al.
Veröffentlicht: (2024)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2024)
Multi-Objective Large Language Model Unlearning
von: Pan, Zibin, et al.
Veröffentlicht: (2024)
von: Pan, Zibin, et al.
Veröffentlicht: (2024)
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
FIT to Forget: Robust Continual Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
von: Spracklen, Joseph, et al.
Veröffentlicht: (2026)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
von: Yang, Yushi, et al.
Veröffentlicht: (2024)
von: Yang, Yushi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
von: Sondej, Filip, et al.
Veröffentlicht: (2025) -
Explainable LLM Unlearning Through Reasoning
von: Liao, Junfeng, et al.
Veröffentlicht: (2026) -
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025) -
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026) -
LLM Unlearning Without an Expert Curated Dataset
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)