LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Chongyu, Wang, Changsheng, Huang, Yancheng, Pal, Soumyadeep, Liu, Sijia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
by: Pal, Soumyadeep, et al.
Published: (2025)
by: Pal, Soumyadeep, et al.
Published: (2025)
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
by: Fan, Chongyu, et al.
Published: (2025)
by: Fan, Chongyu, et al.
Published: (2025)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered
by: Liu, Sijia, et al.
Published: (2026)
by: Liu, Sijia, et al.
Published: (2026)
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
by: Lang, Yicheng, et al.
Published: (2025)
by: Lang, Yicheng, et al.
Published: (2025)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
by: Shang, Bingqi, et al.
Published: (2025)
by: Shang, Bingqi, et al.
Published: (2025)
Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning
by: Jia, Jinghan, et al.
Published: (2024)
by: Jia, Jinghan, et al.
Published: (2024)
Challenging Forgets: Unveiling the Worst-Case Forget Sets in Machine Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
by: Chen, Yiwei, et al.
Published: (2025)
by: Chen, Yiwei, et al.
Published: (2025)
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
by: Reisizadeh, Hadi, et al.
Published: (2025)
by: Reisizadeh, Hadi, et al.
Published: (2025)
LLM Unlearning with LLM Beliefs
by: Li, Kemou, et al.
Published: (2025)
by: Li, Kemou, et al.
Published: (2025)
PrivUn: Unveiling Latent Ripple Effects and Shallow Forgetting in Privacy Unlearning
by: Chen, Xiaoyi, et al.
Published: (2026)
by: Chen, Xiaoyi, et al.
Published: (2026)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
by: Spohn, Philipp, et al.
Published: (2025)
by: Spohn, Philipp, et al.
Published: (2025)
WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
by: Lu, Yuxuan, et al.
Published: (2025)
by: Lu, Yuxuan, et al.
Published: (2025)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
A General Framework to Enhance Fine-tuning-based LLM Unlearning
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
by: Liu, Yezi, et al.
Published: (2025)
by: Liu, Yezi, et al.
Published: (2025)
LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
by: Liu, Yezi, et al.
Published: (2025)
by: Liu, Yezi, et al.
Published: (2025)
Rethinking Machine Unlearning for Large Language Models
by: Liu, Sijia, et al.
Published: (2024)
by: Liu, Sijia, et al.
Published: (2024)
CLUE: Conflict-guided Localization for LLM Unlearning Framework
by: Chen, Hang, et al.
Published: (2025)
by: Chen, Hang, et al.
Published: (2025)
LUME: LLM Unlearning with Multitask Evaluations
by: Ramakrishna, Anil, et al.
Published: (2025)
by: Ramakrishna, Anil, et al.
Published: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024)
by: Zhuang, Haomin, et al.
Published: (2024)
Rotation Control Unlearning: Quantifying and Controlling Continuous Unlearning for LLM with The Cognitive Rotation Space
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Explainable LLM Unlearning Through Reasoning
by: Liao, Junfeng, et al.
Published: (2026)
by: Liao, Junfeng, et al.
Published: (2026)
Label Smoothing Improves Gradient Ascent in LLM Unlearning
by: Pang, Zirui, et al.
Published: (2025)
by: Pang, Zirui, et al.
Published: (2025)
LLM Unlearning via Loss Adjustment with Only Forget Data
by: Wang, Yaxuan, et al.
Published: (2024)
by: Wang, Yaxuan, et al.
Published: (2024)
Soundness-Aware Level: A Microscopic Signature that Predicts LLM Reasoning Potential
by: Wu, Xuansheng, et al.
Published: (2025)
by: Wu, Xuansheng, et al.
Published: (2025)
LLM Unlearning Without an Expert Curated Dataset
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
by: Wang, Yaxuan, et al.
Published: (2025)
by: Wang, Yaxuan, et al.
Published: (2025)
GRU: Mitigating the Trade-off between Unlearning and Retention for LLMs
by: Wang, Yue, et al.
Published: (2025)
by: Wang, Yue, et al.
Published: (2025)
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
by: Qiu, Zhaopeng, et al.
Published: (2026)
by: Qiu, Zhaopeng, et al.
Published: (2026)
Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization
by: Huang, Yancheng, et al.
Published: (2026)
by: Huang, Yancheng, et al.
Published: (2026)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
by: Liu, Zihe, et al.
Published: (2025)
by: Liu, Zihe, et al.
Published: (2025)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
by: Wu, Xiaoyu, et al.
Published: (2025)
by: Wu, Xiaoyu, et al.
Published: (2025)
LIDS: LLM Summary Inference Under the Layered Lens
by: Park, Dylan, et al.
Published: (2026)
by: Park, Dylan, et al.
Published: (2026)
Similar Items
-
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
by: Pal, Soumyadeep, et al.
Published: (2025) -
Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond
by: Fan, Chongyu, et al.
Published: (2025) -
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024) -
Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered
by: Liu, Sijia, et al.
Published: (2026) -
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
by: Lang, Yicheng, et al.
Published: (2025)