Exclusive Unlearning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sasaki, Mutsumi, Nakayama, Kouta, Miyao, Yusuke, Oseki, Yohei, Isonuma, Masaru
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915921472258048
author Sasaki, Mutsumi
Nakayama, Kouta
Miyao, Yusuke
Oseki, Yohei
Isonuma, Masaru
author_facet Sasaki, Mutsumi
Nakayama, Kouta
Miyao, Yusuke
Oseki, Yohei
Isonuma, Masaru
contents When introducing Large Language Models (LLMs) into industrial applications, such as healthcare and education, the risk of generating harmful content becomes a significant challenge. While existing machine unlearning methods can erase specific harmful knowledge and expressions, diverse harmful content makes comprehensive removal difficult. In this study, instead of individually listing targets for forgetting, we propose Exclusive Unlearning (EU), which aims for broad harm removal by extensively forgetting everything except for the knowledge and expressions we wish to retain. We demonstrate that through Exclusive Unlearning, it is possible to obtain a model that ensures safety against a wide range of inputs, including jailbreaks, while maintaining the ability to respond to diverse instructions related to specific domains such as medicine and mathematics.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06154
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Exclusive Unlearning
Sasaki, Mutsumi
Nakayama, Kouta
Miyao, Yusuke
Oseki, Yohei
Isonuma, Masaru
Computation and Language
When introducing Large Language Models (LLMs) into industrial applications, such as healthcare and education, the risk of generating harmful content becomes a significant challenge. While existing machine unlearning methods can erase specific harmful knowledge and expressions, diverse harmful content makes comprehensive removal difficult. In this study, instead of individually listing targets for forgetting, we propose Exclusive Unlearning (EU), which aims for broad harm removal by extensively forgetting everything except for the knowledge and expressions we wish to retain. We demonstrate that through Exclusive Unlearning, it is possible to obtain a model that ensures safety against a wide range of inputs, including jailbreaks, while maintaining the ability to respond to diverse instructions related to specific domains such as medicine and mathematics.
title Exclusive Unlearning
topic Computation and Language
url https://arxiv.org/abs/2604.06154