Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yamashita, Tomoya, Yamanaka, Yuuki, Yamada, Masanori, Miura, Takayuki, Shibahara, Toshiki, Iwata, Tomoharu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911163397177344
author Yamashita, Tomoya
Yamanaka, Yuuki
Yamada, Masanori
Miura, Takayuki
Shibahara, Toshiki
Iwata, Tomoharu
author_facet Yamashita, Tomoya
Yamanaka, Yuuki
Yamada, Masanori
Miura, Takayuki
Shibahara, Toshiki
Iwata, Tomoharu
contents Machine Unlearning (MU) has recently attracted considerable attention as a solution to privacy and copyright issues in large language models (LLMs). Existing MU methods aim to remove specific target sentences from an LLM while minimizing damage to unrelated knowledge. However, these approaches require explicit target sentences and do not support removing broader concepts, such as persons or events. To address this limitation, we introduce Concept Unlearning (CU) as a new requirement for LLM unlearning. We leverage knowledge graphs to represent the LLM's internal knowledge and define CU as removing the forgetting target nodes and associated edges. This graph-based formulation enables a more intuitive unlearning and facilitates the design of more effective methods. We propose a novel method that prompts the LLM to generate knowledge triplets and explanatory sentences about the forgetting target and applies the unlearning process to these representations. Our approach enables more precise and comprehensive concept removal by aligning the unlearning process with the LLM's internal knowledge representations. Experiments on real-world and synthetic datasets demonstrate that our method effectively achieves concept-level unlearning while preserving unrelated knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
Yamashita, Tomoya
Yamanaka, Yuuki
Yamada, Masanori
Miura, Takayuki
Shibahara, Toshiki
Iwata, Tomoharu
Computation and Language
Machine Learning
Machine Unlearning (MU) has recently attracted considerable attention as a solution to privacy and copyright issues in large language models (LLMs). Existing MU methods aim to remove specific target sentences from an LLM while minimizing damage to unrelated knowledge. However, these approaches require explicit target sentences and do not support removing broader concepts, such as persons or events. To address this limitation, we introduce Concept Unlearning (CU) as a new requirement for LLM unlearning. We leverage knowledge graphs to represent the LLM's internal knowledge and define CU as removing the forgetting target nodes and associated edges. This graph-based formulation enables a more intuitive unlearning and facilitates the design of more effective methods. We propose a novel method that prompts the LLM to generate knowledge triplets and explanatory sentences about the forgetting target and applies the unlearning process to these representations. Our approach enables more precise and comprehensive concept removal by aligning the unlearning process with the LLM's internal knowledge representations. Experiments on real-world and synthetic datasets demonstrate that our method effectively achieves concept-level unlearning while preserving unrelated knowledge.
title Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.15621