Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hwang, Kyomin, Kim, Hyeonjin, Kim, Seungyeon, Wee, Sunghyun, Kwak, Nojun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915582089101312
author Hwang, Kyomin
Kim, Hyeonjin
Kim, Seungyeon
Wee, Sunghyun
Kwak, Nojun
author_facet Hwang, Kyomin
Kim, Hyeonjin
Kim, Seungyeon
Wee, Sunghyun
Kwak, Nojun
contents There have been a couple of studies showing that attempting to erase multilingual knowledge using only English data is insufficient for multilingual LLMs. However, their analyses remain highly performance-oriented. In this paper, we switch the point of view to evaluation, and address an additional blind spot which reveals itself when the multilingual LLM is fully finetuned with parallel multilingual dataset before unlearning. Here, language confusion occurs whereby a model responds in language different from that of the input prompt. Language confusion is a problematic phenomenon in unlearning, causing the standard reference-based metrics to fail. We tackle this phenomenon in three steps: (1) introduce N-gram-based Language-Mix (N-Mix) score to quantitatively show the language confusion is pervasive and consistent in multilingual LLMs, (2) demonstrate that reference-based metrics result in false negatives when N-Mix score is high, and(3) suggest the need of new type of unlearning evaluation that can directly assess the content of the generated sentences. We call this type of metrics as semantic-based metric.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
Hwang, Kyomin
Kim, Hyeonjin
Kim, Seungyeon
Wee, Sunghyun
Kwak, Nojun
Computation and Language
Artificial Intelligence
There have been a couple of studies showing that attempting to erase multilingual knowledge using only English data is insufficient for multilingual LLMs. However, their analyses remain highly performance-oriented. In this paper, we switch the point of view to evaluation, and address an additional blind spot which reveals itself when the multilingual LLM is fully finetuned with parallel multilingual dataset before unlearning. Here, language confusion occurs whereby a model responds in language different from that of the input prompt. Language confusion is a problematic phenomenon in unlearning, causing the standard reference-based metrics to fail. We tackle this phenomenon in three steps: (1) introduce N-gram-based Language-Mix (N-Mix) score to quantitatively show the language confusion is pervasive and consistent in multilingual LLMs, (2) demonstrate that reference-based metrics result in false negatives when N-Mix score is high, and(3) suggest the need of new type of unlearning evaluation that can directly assess the content of the generated sentences. We call this type of metrics as semantic-based metric.
title Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.23949