To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Weiss, Christopher, Kreuter, Frauke, Habernal, Ivan
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911811043852288
author Weiss, Christopher
Kreuter, Frauke
Habernal, Ivan
author_facet Weiss, Christopher
Kreuter, Frauke
Habernal, Ivan
contents Although the NLP community has adopted central differential privacy as a go-to framework for privacy-preserving model training or data sharing, the choice and interpretation of the key parameter, privacy budget $\varepsilon$ that governs the strength of privacy protection, remains largely arbitrary. We argue that determining the $\varepsilon$ value should not be solely in the hands of researchers or system developers, but must also take into account the actual people who share their potentially sensitive data. In other words: Would you share your instant messages for $\varepsilon$ of 10? We address this research gap by designing, implementing, and conducting a behavioral experiment (311 lay participants) to study the behavior of people in uncertain decision-making situations with respect to privacy-threatening situations. Framing the risk perception in terms of two realistic NLP scenarios and using a vignette behavioral study help us determine what $\varepsilon$ thresholds would lead lay people to be willing to share sensitive textual data - to our knowledge, the first study of its kind.
format Preprint
id arxiv_https___arxiv_org_abs_2307_06708
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?
Weiss, Christopher
Kreuter, Frauke
Habernal, Ivan
Computation and Language
Cryptography and Security
Although the NLP community has adopted central differential privacy as a go-to framework for privacy-preserving model training or data sharing, the choice and interpretation of the key parameter, privacy budget $\varepsilon$ that governs the strength of privacy protection, remains largely arbitrary. We argue that determining the $\varepsilon$ value should not be solely in the hands of researchers or system developers, but must also take into account the actual people who share their potentially sensitive data. In other words: Would you share your instant messages for $\varepsilon$ of 10? We address this research gap by designing, implementing, and conducting a behavioral experiment (311 lay participants) to study the behavior of people in uncertain decision-making situations with respect to privacy-threatening situations. Framing the risk perception in terms of two realistic NLP scenarios and using a vignette behavioral study help us determine what $\varepsilon$ thresholds would lead lay people to be willing to share sensitive textual data - to our knowledge, the first study of its kind.
title To share or not to share: What risks would laypeople accept to give sensitive data to differentially-private NLP systems?
topic Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2307.06708