Towards the Anonymization of the Language Modeling

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Boutet, Antoine, Magnana, Lucas, Sénéchal, Juliette
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914582932488192
author Boutet, Antoine
Magnana, Lucas
Sénéchal, Juliette
author_facet Boutet, Antoine
Magnana, Lucas
Sénéchal, Juliette
contents Rapid advances in Natural Language Processing (NLP) have revolutionized many fields, including healthcare. However, these advances raise significant privacy concerns, especially when pre-trained models fine-tuned and specialized on sensitive data can memorize and then expose and regurgitate personal information. This paper presents a privacy-preserving language modeling approach to address the problem of language models anonymization, and thus promote their sharing. Specifically, we propose both a Masking Language Modeling (MLM) methodology to specialize a BERT-like language model, and a Causal Language Modeling (CLM) methodology to specialize a GPT-like model that avoids the model from memorizing direct and indirect identifying information present in the training data. We have comprehensively evaluated our approaches using a medical dataset and compared them against different baselines. Our results indicate that by avoiding memorizing both direct and indirect identifiers during model specialization, our masking and causal language modeling schemes offer a good tradeoff for maintaining high privacy while retaining high utility.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02407
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards the Anonymization of the Language Modeling
Boutet, Antoine
Magnana, Lucas
Sénéchal, Juliette
Computation and Language
Cryptography and Security
Machine Learning
Rapid advances in Natural Language Processing (NLP) have revolutionized many fields, including healthcare. However, these advances raise significant privacy concerns, especially when pre-trained models fine-tuned and specialized on sensitive data can memorize and then expose and regurgitate personal information. This paper presents a privacy-preserving language modeling approach to address the problem of language models anonymization, and thus promote their sharing. Specifically, we propose both a Masking Language Modeling (MLM) methodology to specialize a BERT-like language model, and a Causal Language Modeling (CLM) methodology to specialize a GPT-like model that avoids the model from memorizing direct and indirect identifying information present in the training data. We have comprehensively evaluated our approaches using a medical dataset and compared them against different baselines. Our results indicate that by avoiding memorizing both direct and indirect identifiers during model specialization, our masking and causal language modeling schemes offer a good tradeoff for maintaining high privacy while retaining high utility.
title Towards the Anonymization of the Language Modeling
topic Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2501.02407