Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kadhe, Swanand Ravindra, Ahmed, Farhan, Wei, Dennis, Baracaldo, Nathalie, Padhi, Inkit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating the Dynamics of Membership Privacy in Deep Learning
von: Chen, Yuetian, et al.
Veröffentlicht: (2025)
von: Chen, Yuetian, et al.
Veröffentlicht: (2025)
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
von: Djuhera, Aladin, et al.
Veröffentlicht: (2026)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2026)
SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
von: Jiang, Shuli, et al.
Veröffentlicht: (2024)
von: Jiang, Shuli, et al.
Veröffentlicht: (2024)
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
How Data Inter-connectivity Shapes LLMs Unlearning: A Structural Unlearning Perspective
von: Qiu, Xinchi, et al.
Veröffentlicht: (2024)
von: Qiu, Xinchi, et al.
Veröffentlicht: (2024)
STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
von: An, Sungeun, et al.
Veröffentlicht: (2026)
von: An, Sungeun, et al.
Veröffentlicht: (2026)
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
von: Niu, Peizhi, et al.
Veröffentlicht: (2025)
von: Niu, Peizhi, et al.
Veröffentlicht: (2025)
Tool Unlearning for Tool-Augmented LLMs
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
von: Wang, Changsheng, et al.
Veröffentlicht: (2025)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
von: Hou, Yufang, et al.
Veröffentlicht: (2024)
von: Hou, Yufang, et al.
Veröffentlicht: (2024)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
LLM Unlearning via Loss Adjustment with Only Forget Data
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
Geometric-disentangelment Unlearning
von: Zhou, Duo, et al.
Veröffentlicht: (2025)
von: Zhou, Duo, et al.
Veröffentlicht: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
von: Sondej, Filip, et al.
Veröffentlicht: (2025)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2024)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
von: Sinha, Yash, et al.
Veröffentlicht: (2024)
von: Sinha, Yash, et al.
Veröffentlicht: (2024)
Textual Unlearning Gives a False Sense of Unlearning
von: Du, Jiacheng, et al.
Veröffentlicht: (2024)
von: Du, Jiacheng, et al.
Veröffentlicht: (2024)
Large Language Model Unlearning
von: Yao, Yuanshun, et al.
Veröffentlicht: (2023)
von: Yao, Yuanshun, et al.
Veröffentlicht: (2023)
In-Context Probing for Membership Inference in Fine-Tuned Language Models
von: Lu, Zhexi, et al.
Veröffentlicht: (2025)
von: Lu, Zhexi, et al.
Veröffentlicht: (2025)
NegMerge: Sign-Consensual Weight Merging for Machine Unlearning
von: Kim, Hyo Seo, et al.
Veröffentlicht: (2024)
von: Kim, Hyo Seo, et al.
Veröffentlicht: (2024)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
von: Xie, Linxi, et al.
Veröffentlicht: (2025)
von: Xie, Linxi, et al.
Veröffentlicht: (2025)
A Neuro-inspired Interpretation of Unlearning in Large Language Models through Sample-level Unlearning Difficulty
von: Feng, Xiaohua, et al.
Veröffentlicht: (2025)
von: Feng, Xiaohua, et al.
Veröffentlicht: (2025)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
EvoMU: Evolutionary Machine Unlearning
von: Batorski, Pawel, et al.
Veröffentlicht: (2026)
von: Batorski, Pawel, et al.
Veröffentlicht: (2026)
Explainable LLM Unlearning Through Reasoning
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
von: Liao, Junfeng, et al.
Veröffentlicht: (2026)
Offset Unlearning for Large Language Models
von: Huang, James Y., et al.
Veröffentlicht: (2024)
von: Huang, James Y., et al.
Veröffentlicht: (2024)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
von: Shumailov, Ilia, et al.
Veröffentlicht: (2024)
von: Shumailov, Ilia, et al.
Veröffentlicht: (2024)
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
von: Mahmud, Tamim Al, et al.
Veröffentlicht: (2025)
von: Mahmud, Tamim Al, et al.
Veröffentlicht: (2025)
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yaxuan, et al.
Veröffentlicht: (2025)
Leverage Unlearning to Sanitize LLMs
von: Boutet, Antoine, et al.
Veröffentlicht: (2025)
von: Boutet, Antoine, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating the Dynamics of Membership Privacy in Deep Learning
von: Chen, Yuetian, et al.
Veröffentlicht: (2025) -
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
von: Djuhera, Aladin, et al.
Veröffentlicht: (2026) -
SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025) -
Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
von: Jiang, Shuli, et al.
Veröffentlicht: (2024) -
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)