LLM Unlearning Without an Expert Curated Dataset

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Xiaoyuan, Zhang, Muru, Liu, Ollie, Jia, Robin, Neiswanger, Willie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908578638462976
author Zhu, Xiaoyuan
Zhang, Muru
Liu, Ollie
Jia, Robin
Neiswanger, Willie
author_facet Zhu, Xiaoyuan
Zhang, Muru
Liu, Ollie
Jia, Robin
Neiswanger, Willie
contents Modern large language models often encode sensitive, harmful, or copyrighted knowledge, raising the need for post-hoc unlearning-the ability to remove specific domains of knowledge from a model without full retraining. A major bottleneck in current unlearning pipelines is constructing effective forget sets-datasets that approximate the target domain and guide the model to forget it. In this work, we introduce a scalable, automated approach to generate high-quality forget sets using language models themselves. Our method synthesizes textbook-style data through a structured prompting pipeline, requiring only a domain name as input. Through experiments on unlearning biosecurity, cybersecurity, and Harry Potter novels, we show that our synthetic datasets consistently outperform the baseline synthetic alternatives and are comparable to the expert-curated ones. Additionally, ablation studies reveal that the multi-step generation pipeline significantly boosts data diversity, which in turn improves unlearning utility. Overall, our findings suggest that synthetic datasets offer a promising path toward practical, scalable unlearning for a wide range of emerging domains without the need for manual intervention. We release our code and dataset at https://github.com/xyzhu123/Synthetic_Textbook.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06595
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Unlearning Without an Expert Curated Dataset
Zhu, Xiaoyuan
Zhang, Muru
Liu, Ollie
Jia, Robin
Neiswanger, Willie
Computation and Language
Artificial Intelligence
Machine Learning
Modern large language models often encode sensitive, harmful, or copyrighted knowledge, raising the need for post-hoc unlearning-the ability to remove specific domains of knowledge from a model without full retraining. A major bottleneck in current unlearning pipelines is constructing effective forget sets-datasets that approximate the target domain and guide the model to forget it. In this work, we introduce a scalable, automated approach to generate high-quality forget sets using language models themselves. Our method synthesizes textbook-style data through a structured prompting pipeline, requiring only a domain name as input. Through experiments on unlearning biosecurity, cybersecurity, and Harry Potter novels, we show that our synthetic datasets consistently outperform the baseline synthetic alternatives and are comparable to the expert-curated ones. Additionally, ablation studies reveal that the multi-step generation pipeline significantly boosts data diversity, which in turn improves unlearning utility. Overall, our findings suggest that synthetic datasets offer a promising path toward practical, scalable unlearning for a wide range of emerging domains without the need for manual intervention. We release our code and dataset at https://github.com/xyzhu123/Synthetic_Textbook.
title LLM Unlearning Without an Expert Curated Dataset
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.06595