I Prefer not to Say: Protecting User Consent in Models with Optional Personal Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Leemann, Tobias, Pawelczyk, Martin, Eberle, Christian Thomas, Kasneci, Gjergji
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910314999578624
author Leemann, Tobias
Pawelczyk, Martin
Eberle, Christian Thomas
Kasneci, Gjergji
author_facet Leemann, Tobias
Pawelczyk, Martin
Eberle, Christian Thomas
Kasneci, Gjergji
contents We examine machine learning models in a setup where individuals have the choice to share optional personal information with a decision-making system, as seen in modern insurance pricing models. Some users consent to their data being used whereas others object and keep their data undisclosed. In this work, we show that the decision not to share data can be considered as information in itself that should be protected to respect users' privacy. This observation raises the overlooked problem of how to ensure that users who protect their personal data do not suffer any disadvantages as a result. To address this problem, we formalize protection requirements for models which only use the information for which active user consent was obtained. This excludes implicit information contained in the decision to share data or not. We offer the first solution to this problem by proposing the notion of Protected User Consent (PUC), which we prove to be loss-optimal under our protection requirement. We observe that privacy and performance are not fundamentally at odds with each other and that it is possible for a decision maker to benefit from additional data while respecting users' consent. To learn PUC-compliant models, we devise a model-agnostic data augmentation strategy with finite sample convergence guarantees. Finally, we analyze the implications of PUC on challenging real datasets, tasks, and models.
format Preprint
id arxiv_https___arxiv_org_abs_2210_13954
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle I Prefer not to Say: Protecting User Consent in Models with Optional Personal Data
Leemann, Tobias
Pawelczyk, Martin
Eberle, Christian Thomas
Kasneci, Gjergji
Machine Learning
Artificial Intelligence
Computers and Society
We examine machine learning models in a setup where individuals have the choice to share optional personal information with a decision-making system, as seen in modern insurance pricing models. Some users consent to their data being used whereas others object and keep their data undisclosed. In this work, we show that the decision not to share data can be considered as information in itself that should be protected to respect users' privacy. This observation raises the overlooked problem of how to ensure that users who protect their personal data do not suffer any disadvantages as a result. To address this problem, we formalize protection requirements for models which only use the information for which active user consent was obtained. This excludes implicit information contained in the decision to share data or not. We offer the first solution to this problem by proposing the notion of Protected User Consent (PUC), which we prove to be loss-optimal under our protection requirement. We observe that privacy and performance are not fundamentally at odds with each other and that it is possible for a decision maker to benefit from additional data while respecting users' consent. To learn PUC-compliant models, we devise a model-agnostic data augmentation strategy with finite sample convergence guarantees. Finally, we analyze the implications of PUC on challenging real datasets, tasks, and models.
title I Prefer not to Say: Protecting User Consent in Models with Optional Personal Data
topic Machine Learning
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2210.13954