Exclusive Unlearning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915921472258048 |
|---|---|
| author | Sasaki, Mutsumi Nakayama, Kouta Miyao, Yusuke Oseki, Yohei Isonuma, Masaru |
| author_facet | Sasaki, Mutsumi Nakayama, Kouta Miyao, Yusuke Oseki, Yohei Isonuma, Masaru |
| contents | When introducing Large Language Models (LLMs) into industrial applications, such as healthcare and education, the risk of generating harmful content becomes a significant challenge. While existing machine unlearning methods can erase specific harmful knowledge and expressions, diverse harmful content makes comprehensive removal difficult. In this study, instead of individually listing targets for forgetting, we propose Exclusive Unlearning (EU), which aims for broad harm removal by extensively forgetting everything except for the knowledge and expressions we wish to retain. We demonstrate that through Exclusive Unlearning, it is possible to obtain a model that ensures safety against a wide range of inputs, including jailbreaks, while maintaining the ability to respond to diverse instructions related to specific domains such as medicine and mathematics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_06154 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Exclusive Unlearning Sasaki, Mutsumi Nakayama, Kouta Miyao, Yusuke Oseki, Yohei Isonuma, Masaru Computation and Language When introducing Large Language Models (LLMs) into industrial applications, such as healthcare and education, the risk of generating harmful content becomes a significant challenge. While existing machine unlearning methods can erase specific harmful knowledge and expressions, diverse harmful content makes comprehensive removal difficult. In this study, instead of individually listing targets for forgetting, we propose Exclusive Unlearning (EU), which aims for broad harm removal by extensively forgetting everything except for the knowledge and expressions we wish to retain. We demonstrate that through Exclusive Unlearning, it is possible to obtain a model that ensures safety against a wide range of inputs, including jailbreaks, while maintaining the ability to respond to diverse instructions related to specific domains such as medicine and mathematics. |
| title | Exclusive Unlearning |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2604.06154 |