The Unlearning Mirage: A Dynamic Framework for Evaluating LLM Unlearning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shah, Raj Sanjay, Huang, Jing, Murugesan, Keerthiram, Baracaldo, Nathalie, Yang, Diyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024)
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
Towards Evaluation for Real-World LLM Unlearning
von: Miao, Ke, et al.
Veröffentlicht: (2025)
von: Miao, Ke, et al.
Veröffentlicht: (2025)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
von: Wang, Bin, et al.
Veröffentlicht: (2026)
von: Wang, Bin, et al.
Veröffentlicht: (2026)
Machine Unlearning under Overparameterization
von: Block, Jacob L., et al.
Veröffentlicht: (2025)
von: Block, Jacob L., et al.
Veröffentlicht: (2025)
Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning
von: Yu, Zhenyu, et al.
Veröffentlicht: (2026)
von: Yu, Zhenyu, et al.
Veröffentlicht: (2026)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
von: Mahmud, Tamim Al, et al.
Veröffentlicht: (2025)
von: Mahmud, Tamim Al, et al.
Veröffentlicht: (2025)
STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language Models
von: Basavatia, Shreyas, et al.
Veröffentlicht: (2024)
von: Basavatia, Shreyas, et al.
Veröffentlicht: (2024)
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
von: Yang, Puning, et al.
Veröffentlicht: (2025)
von: Yang, Puning, et al.
Veröffentlicht: (2025)
Machine Unlearning with Minimal Gradient Dependence for High Unlearning Ratios
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
von: Chen, Yiwei, et al.
Veröffentlicht: (2025)
von: Chen, Yiwei, et al.
Veröffentlicht: (2025)
A Reliable Cryptographic Framework for Empirical Machine Unlearning Evaluation
von: Tu, Yiwen, et al.
Veröffentlicht: (2024)
von: Tu, Yiwen, et al.
Veröffentlicht: (2024)
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
Unlearning Information Bottleneck: Machine Unlearning of Systematic Patterns and Biases
von: Han, Ling, et al.
Veröffentlicht: (2024)
von: Han, Ling, et al.
Veröffentlicht: (2024)
MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts
von: Gu, Tianle, et al.
Veröffentlicht: (2024)
von: Gu, Tianle, et al.
Veröffentlicht: (2024)
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
Unlearning via Sparse Representations
von: Shah, Vedant, et al.
Veröffentlicht: (2023)
von: Shah, Vedant, et al.
Veröffentlicht: (2023)
Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation of LLM Unlearning Robustness under Multi-Turn Interaction
von: Pan, Ruihao, et al.
Veröffentlicht: (2026)
von: Pan, Ruihao, et al.
Veröffentlicht: (2026)
Learn to Unlearn: Meta-Learning-Based Knowledge Graph Embedding Unlearning
von: Xu, Naixing, et al.
Veröffentlicht: (2024)
von: Xu, Naixing, et al.
Veröffentlicht: (2024)
Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit Difference
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
von: Ji, Jiabao, et al.
Veröffentlicht: (2024)
Evaluating the Dynamics of Membership Privacy in Deep Learning
von: Chen, Yuetian, et al.
Veröffentlicht: (2025)
von: Chen, Yuetian, et al.
Veröffentlicht: (2025)
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
von: Zhong, Yisheng, et al.
Veröffentlicht: (2026)
von: Zhong, Yisheng, et al.
Veröffentlicht: (2026)
Deep Unlearn: Benchmarking Machine Unlearning for Image Classification
von: Cadet, Xavier F., et al.
Veröffentlicht: (2024)
von: Cadet, Xavier F., et al.
Veröffentlicht: (2024)
Context Attribution with Multi-Armed Bandit Optimization
von: Pan, Deng, et al.
Veröffentlicht: (2025)
von: Pan, Deng, et al.
Veröffentlicht: (2025)
LLM Unlearning via Neural Activation Redirection
von: Shen, William F., et al.
Veröffentlicht: (2025)
von: Shen, William F., et al.
Veröffentlicht: (2025)
Agents Are All You Need for LLM Unlearning
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
Catastrophic Failure of LLM Unlearning via Quantization
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
von: Villa, Danielle, et al.
Veröffentlicht: (2025)
von: Villa, Danielle, et al.
Veröffentlicht: (2025)
Explainable LLM Unlearning Through Reasoning
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
Interference-Aware Multi-Task Unlearning
von: Huang, Ying-Hua, et al.
Veröffentlicht: (2026)
von: Huang, Ying-Hua, et al.
Veröffentlicht: (2026)
Ready2Unlearn: A Learning-Time Approach for Preparing Models with Future Unlearning Readiness
von: Duan, Hanyu, et al.
Veröffentlicht: (2025)
von: Duan, Hanyu, et al.
Veröffentlicht: (2025)
Unlink to Unlearn: Simplifying Edge Unlearning in GNNs
von: Tan, Jiajun, et al.
Veröffentlicht: (2024)
von: Tan, Jiajun, et al.
Veröffentlicht: (2024)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework
von: Shah, Aakriti, et al.
Veröffentlicht: (2025)
von: Shah, Aakriti, et al.
Veröffentlicht: (2025)
Safety Alignment via Constrained Knowledge Unlearning
von: Shi, Zesheng, et al.
Veröffentlicht: (2025)
von: Shi, Zesheng, et al.
Veröffentlicht: (2025)
An Information Theoretic Evaluation Metric For Strong Unlearning
von: Jeon, Dongjae, et al.
Veröffentlicht: (2024)
von: Jeon, Dongjae, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
von: Kadhe, Swanand Ravindra, et al.
Veröffentlicht: (2024) -
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
von: Wang, Changsheng, et al.
Veröffentlicht: (2025) -
Towards Evaluation for Real-World LLM Unlearning
von: Miao, Ke, et al.
Veröffentlicht: (2025) -
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
von: Wang, Bin, et al.
Veröffentlicht: (2026) -
Machine Unlearning under Overparameterization
von: Block, Jacob L., et al.
Veröffentlicht: (2025)