IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Frikha, Ahmed, Walha, Nassim, Nakka, Krishna Kanth, Mendes, Ricardo, Jiang, Xue, Zhou, Xuebing
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915133875290112
author Frikha, Ahmed
Walha, Nassim
Nakka, Krishna Kanth
Mendes, Ricardo
Jiang, Xue
Zhou, Xuebing
author_facet Frikha, Ahmed
Walha, Nassim
Nakka, Krishna Kanth
Mendes, Ricardo
Jiang, Xue
Zhou, Xuebing
contents In this work, we address the problem of text anonymization where the goal is to prevent adversaries from correctly inferring private attributes of the author, while keeping the text utility, i.e., meaning and semantics. We propose IncogniText, a technique that anonymizes the text to mislead a potential adversary into predicting a wrong private attribute value. Our empirical evaluation shows a reduction of private attribute leakage by more than 90% across 8 different private attributes. Finally, we demonstrate the maturity of IncogniText for real-world applications by distilling its anonymization capability into a set of LoRA parameters associated with an on-device model. Our results show the possibility of reducing privacy leakage by more than half with limited impact on utility.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02956
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization
Frikha, Ahmed
Walha, Nassim
Nakka, Krishna Kanth
Mendes, Ricardo
Jiang, Xue
Zhou, Xuebing
Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
In this work, we address the problem of text anonymization where the goal is to prevent adversaries from correctly inferring private attributes of the author, while keeping the text utility, i.e., meaning and semantics. We propose IncogniText, a technique that anonymizes the text to mislead a potential adversary into predicting a wrong private attribute value. Our empirical evaluation shows a reduction of private attribute leakage by more than 90% across 8 different private attributes. Finally, we demonstrate the maturity of IncogniText for real-world applications by distilling its anonymization capability into a set of LoRA parameters associated with an on-device model. Our results show the possibility of reducing privacy leakage by more than half with limited impact on utility.
title IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.02956