Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Scholten, Yan, Xhonneux, Sophie, Schwinn, Leo, Günnemann, Stephan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
by: Scholten, Yan, et al.
Published: (2024)
by: Scholten, Yan, et al.
Published: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
by: Seyitoğlu, Atakan, et al.
Published: (2024)
by: Seyitoğlu, Atakan, et al.
Published: (2024)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
by: Schwinn, Leo, et al.
Published: (2024)
by: Schwinn, Leo, et al.
Published: (2024)
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
by: Schwinn, Leo, et al.
Published: (2025)
by: Schwinn, Leo, et al.
Published: (2025)
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Assessing Robustness via Score-Based Adversarial Image Generation
by: Kollovieh, Marcel, et al.
Published: (2023)
by: Kollovieh, Marcel, et al.
Published: (2023)
Sampling-aware Adversarial Attacks Against Large Language Models
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
Flow Matching with Gaussian Process Priors for Probabilistic Time Series Forecasting
by: Kollovieh, Marcel, et al.
Published: (2024)
by: Kollovieh, Marcel, et al.
Published: (2024)
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Hierarchical Randomized Smoothing
by: Scholten, Yan, et al.
Published: (2023)
by: Scholten, Yan, et al.
Published: (2023)
Provably Reliable Conformal Prediction Sets in the Presence of Data Poisoning
by: Scholten, Yan, et al.
Published: (2024)
by: Scholten, Yan, et al.
Published: (2024)
Efficient Time Series Processing for Transformers and State-Space Models through Token Merging
by: Götz, Leon, et al.
Published: (2024)
by: Götz, Leon, et al.
Published: (2024)
Diffusion LLMs are Natural Adversaries for any LLM
by: Lüdke, David, et al.
Published: (2025)
by: Lüdke, David, et al.
Published: (2025)
Predictive Feature Caching for Training-free Acceleration of Molecular Geometry Generation
by: Sommer, Johanna, et al.
Published: (2025)
by: Sommer, Johanna, et al.
Published: (2025)
Feature-Selective Representation Misdirection for Machine Unlearning
by: Chen, Taozhao, et al.
Published: (2025)
by: Chen, Taozhao, et al.
Published: (2025)
Machine Unlearning in Low-Dimensional Feature Subspace
by: Fang, Kun, et al.
Published: (2026)
by: Fang, Kun, et al.
Published: (2026)
Joint Relational Database Generation via Graph-Conditional Diffusion Models
by: Ketata, Mohamed Amine, et al.
Published: (2025)
by: Ketata, Mohamed Amine, et al.
Published: (2025)
Byte Pair Encoding for Efficient Time Series Forecasting
by: Götz, Leon, et al.
Published: (2025)
by: Götz, Leon, et al.
Published: (2025)
Fast Proxies for LLM Robustness Evaluation
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Intriguing Properties of Input-dependent Randomized Smoothing
by: Súkeník, Peter, et al.
Published: (2021)
by: Súkeník, Peter, et al.
Published: (2021)
Automated Machine Learning: A Case Study on Non-Intrusive Appliance Load Monitoring
by: Moin, Armin, et al.
Published: (2022)
by: Moin, Armin, et al.
Published: (2022)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
Joint Out-of-Distribution Filtering and Data Discovery Active Learning
by: Schmidt, Sebastian, et al.
Published: (2025)
by: Schmidt, Sebastian, et al.
Published: (2025)
Long-Range Graph Wavelet Networks
by: Guerranti, Filippo, et al.
Published: (2025)
by: Guerranti, Filippo, et al.
Published: (2025)
Spatio-Spectral Graph Neural Networks
by: Geisler, Simon, et al.
Published: (2024)
by: Geisler, Simon, et al.
Published: (2024)
Deep Unlearn: Benchmarking Machine Unlearning for Image Classification
by: Cadet, Xavier F., et al.
Published: (2024)
by: Cadet, Xavier F., et al.
Published: (2024)
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
by: Mahmud, Tamim Al, et al.
Published: (2025)
by: Mahmud, Tamim Al, et al.
Published: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
Unlearning Information Bottleneck: Machine Unlearning of Systematic Patterns and Biases
by: Han, Ling, et al.
Published: (2024)
by: Han, Ling, et al.
Published: (2024)
Adversarial Robustness of Graph Transformers
by: Foth, Philipp, et al.
Published: (2024)
by: Foth, Philipp, et al.
Published: (2024)
Auditing Approximate Machine Unlearning for Differentially Private Models
by: Gu, Yuechun, et al.
Published: (2025)
by: Gu, Yuechun, et al.
Published: (2025)
Machine Unlearning under Overparameterization
by: Block, Jacob L., et al.
Published: (2025)
by: Block, Jacob L., et al.
Published: (2025)
Group-robust Machine Unlearning
by: De Min, Thomas, et al.
Published: (2025)
by: De Min, Thomas, et al.
Published: (2025)
Soft Weighted Machine Unlearning
by: Qiao, Xinbao, et al.
Published: (2025)
by: Qiao, Xinbao, et al.
Published: (2025)
Machine Unlearning: Solutions and Challenges
by: Xu, Jie, et al.
Published: (2023)
by: Xu, Jie, et al.
Published: (2023)
A Survey of Machine Unlearning
by: Nguyen, Thanh Tam, et al.
Published: (2022)
by: Nguyen, Thanh Tam, et al.
Published: (2022)
Machine Unlearning in Contrastive Learning
by: Wang, Zixin, et al.
Published: (2024)
by: Wang, Zixin, et al.
Published: (2024)
Similar Items
-
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
by: Scholten, Yan, et al.
Published: (2024) -
Extracting Unlearned Information from LLMs with Activation Steering
by: Seyitoğlu, Atakan, et al.
Published: (2024) -
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
by: Schwinn, Leo, et al.
Published: (2024) -
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
by: Schwinn, Leo, et al.
Published: (2025) -
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)