Cross-Lingual Unlearning of Selective Knowledge in Multilingual Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Minseok, Min, Kyunghyun, Choo, Jaegul
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914963213254656
author Choi, Minseok
Min, Kyunghyun
Choo, Jaegul
author_facet Choi, Minseok
Min, Kyunghyun
Choo, Jaegul
contents Pretrained language models memorize vast amounts of information, including private and copyrighted data, raising significant safety concerns. Retraining these models after excluding sensitive data is prohibitively expensive, making machine unlearning a viable, cost-effective alternative. Previous research has focused on machine unlearning for monolingual models, but we find that unlearning in one language does not necessarily transfer to others. This vulnerability makes models susceptible to low-resource language attacks, where sensitive information remains accessible in less dominant languages. This paper presents a pioneering approach to machine unlearning for multilingual language models, selectively erasing information across different languages while maintaining overall performance. Specifically, our method employs an adaptive unlearning scheme that assigns language-dependent weights to address different language performances of multilingual language models. Empirical results demonstrate the effectiveness of our framework compared to existing unlearning baselines, setting a new standard for secure and adaptable multilingual language models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12354
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Lingual Unlearning of Selective Knowledge in Multilingual Language Models
Choi, Minseok
Min, Kyunghyun
Choo, Jaegul
Computation and Language
Pretrained language models memorize vast amounts of information, including private and copyrighted data, raising significant safety concerns. Retraining these models after excluding sensitive data is prohibitively expensive, making machine unlearning a viable, cost-effective alternative. Previous research has focused on machine unlearning for monolingual models, but we find that unlearning in one language does not necessarily transfer to others. This vulnerability makes models susceptible to low-resource language attacks, where sensitive information remains accessible in less dominant languages. This paper presents a pioneering approach to machine unlearning for multilingual language models, selectively erasing information across different languages while maintaining overall performance. Specifically, our method employs an adaptive unlearning scheme that assigns language-dependent weights to address different language performances of multilingual language models. Empirical results demonstrate the effectiveness of our framework compared to existing unlearning baselines, setting a new standard for secure and adaptable multilingual language models.
title Cross-Lingual Unlearning of Selective Knowledge in Multilingual Language Models
topic Computation and Language
url https://arxiv.org/abs/2406.12354