Keep It Private: Unsupervised Privatization of Online Text

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bao, Calvin, Carpuat, Marine
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929346119204864
author Bao, Calvin
Carpuat, Marine
author_facet Bao, Calvin
Carpuat, Marine
contents Authorship obfuscation techniques hold the promise of helping people protect their privacy in online communications by automatically rewriting text to hide the identity of the original author. However, obfuscation has been evaluated in narrow settings in the NLP literature and has primarily been addressed with superficial edit operations that can lead to unnatural outputs. In this work, we introduce an automatic text privatization framework that fine-tunes a large language model via reinforcement learning to produce rewrites that balance soundness, sense, and privacy. We evaluate it extensively on a large-scale test set of English Reddit posts by 68k authors composed of short-medium length texts. We study how the performance changes among evaluative conditions including authorial profile length and authorship detection strategy. Our method maintains high text quality according to both automated metrics and human evaluation, and successfully evades several automated authorship attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10260
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Keep It Private: Unsupervised Privatization of Online Text
Bao, Calvin
Carpuat, Marine
Computation and Language
Artificial Intelligence
Authorship obfuscation techniques hold the promise of helping people protect their privacy in online communications by automatically rewriting text to hide the identity of the original author. However, obfuscation has been evaluated in narrow settings in the NLP literature and has primarily been addressed with superficial edit operations that can lead to unnatural outputs. In this work, we introduce an automatic text privatization framework that fine-tunes a large language model via reinforcement learning to produce rewrites that balance soundness, sense, and privacy. We evaluate it extensively on a large-scale test set of English Reddit posts by 68k authors composed of short-medium length texts. We study how the performance changes among evaluative conditions including authorial profile length and authorship detection strategy. Our method maintains high text quality according to both automated metrics and human evaluation, and successfully evades several automated authorship attacks.
title Keep It Private: Unsupervised Privatization of Online Text
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.10260