Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yiwei, Pal, Soumyadeep, Zhang, Yimeng, Qu, Qing, Liu, Sijia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
by: Reisizadeh, Hadi, et al.
Published: (2025)
by: Reisizadeh, Hadi, et al.
Published: (2025)
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
by: Fan, Chongyu, et al.
Published: (2025)
by: Fan, Chongyu, et al.
Published: (2025)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
by: Pal, Soumyadeep, et al.
Published: (2025)
by: Pal, Soumyadeep, et al.
Published: (2025)
Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
by: Hu, Shengyuan, et al.
Published: (2024)
by: Hu, Shengyuan, et al.
Published: (2024)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
by: Shang, Bingqi, et al.
Published: (2025)
by: Shang, Bingqi, et al.
Published: (2025)
SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning
by: Jia, Jinghan, et al.
Published: (2024)
by: Jia, Jinghan, et al.
Published: (2024)
When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
by: Galeone, Cosimo, et al.
Published: (2026)
by: Galeone, Cosimo, et al.
Published: (2026)
The Dual Power of Interpretable Token Embeddings: Jailbreaking Attacks and Defenses for Diffusion Model Unlearning
by: Chen, Siyi, et al.
Published: (2025)
by: Chen, Siyi, et al.
Published: (2025)
ML Interpretability: Simple Isn't Easy
by: Räz, Tim
Published: (2022)
by: Räz, Tim
Published: (2022)
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
by: Chen, Yiwei, et al.
Published: (2025)
by: Chen, Yiwei, et al.
Published: (2025)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
by: Zhuang, Haomin, et al.
Published: (2024)
by: Zhuang, Haomin, et al.
Published: (2024)
Model Sparsity Can Simplify Machine Unlearning
by: Jia, Jinghan, et al.
Published: (2023)
by: Jia, Jinghan, et al.
Published: (2023)
Explainable AI Isn't Enough! Rethinking Algorithmic Contestability
by: Freiesleben, Timo, et al.
Published: (2026)
by: Freiesleben, Timo, et al.
Published: (2026)
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
by: Favero, Alessandro, et al.
Published: (2025)
by: Favero, Alessandro, et al.
Published: (2025)
WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models
by: Jia, Jinghan, et al.
Published: (2024)
by: Jia, Jinghan, et al.
Published: (2024)
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
Label Smoothing Improves Machine Unlearning
by: Di, Zonglin, et al.
Published: (2024)
by: Di, Zonglin, et al.
Published: (2024)
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
by: Lang, Yicheng, et al.
Published: (2025)
by: Lang, Yicheng, et al.
Published: (2025)
Why Isn't Relational Learning Taking Over the World?
by: Poole, David
Published: (2025)
by: Poole, David
Published: (2025)
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
by: Mahmud, Tamim Al, et al.
Published: (2025)
by: Mahmud, Tamim Al, et al.
Published: (2025)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
by: Tang, Haoyu, et al.
Published: (2024)
by: Tang, Haoyu, et al.
Published: (2024)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models
by: Mustaqim, S. M., et al.
Published: (2025)
by: Mustaqim, S. M., et al.
Published: (2025)
Causal Inference Isn't Special: Why It's Just Another Prediction Problem
by: Fernández-Loría, Carlos
Published: (2025)
by: Fernández-Loría, Carlos
Published: (2025)
On the Necessity of Output Distribution Reweighting for Effective Class Unlearning
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
by: Ebrahimpour-Boroojeny, Ali, et al.
Published: (2025)
Leverage Unlearning to Sanitize LLMs
by: Boutet, Antoine, et al.
Published: (2025)
by: Boutet, Antoine, et al.
Published: (2025)
Rethinking Machine Unlearning for Large Language Models
by: Liu, Sijia, et al.
Published: (2024)
by: Liu, Sijia, et al.
Published: (2024)
Unified Parameter-Efficient Unlearning for LLMs
by: Ding, Chenlu, et al.
Published: (2024)
by: Ding, Chenlu, et al.
Published: (2024)
Harmonizing Multi-Objective LLM Unlearning via Unified Domain Representation and Bidirectional Logit Distillation
by: Zhong, Yisheng, et al.
Published: (2026)
by: Zhong, Yisheng, et al.
Published: (2026)
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
by: Li, Xiangman, et al.
Published: (2025)
by: Li, Xiangman, et al.
Published: (2025)
Targeted Unlearning with Single Layer Unlearning Gradient
by: Cai, Zikui, et al.
Published: (2024)
by: Cai, Zikui, et al.
Published: (2024)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
by: Qiu, Xinchi, et al.
Published: (2024)
by: Qiu, Xinchi, et al.
Published: (2024)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning
by: Barale, Claire, et al.
Published: (2025)
by: Barale, Claire, et al.
Published: (2025)
Probing Knowledge Holes in Unlearned LLMs
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
Parallel Unlearning in Inherited Model Networks
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Similar Items
-
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025) -
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
by: Reisizadeh, Hadi, et al.
Published: (2025) -
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
by: Fan, Chongyu, et al.
Published: (2025) -
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
by: Pal, Soumyadeep, et al.
Published: (2025) -
Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
by: Wang, Changsheng, et al.
Published: (2025)