Preserving Expert-Level Privacy in Offline Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sharma, Navodita, Vinod, Vishnu, Thakurta, Abhradeep, Agarwal, Alekh, Balle, Borja, Dann, Christoph, Raghuveer, Aravindan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918215173537792
author Sharma, Navodita
Vinod, Vishnu
Thakurta, Abhradeep
Agarwal, Alekh
Balle, Borja
Dann, Christoph
Raghuveer, Aravindan
author_facet Sharma, Navodita
Vinod, Vishnu
Thakurta, Abhradeep
Agarwal, Alekh
Balle, Borja
Dann, Christoph
Raghuveer, Aravindan
contents The offline reinforcement learning (RL) problem aims to learn an optimal policy from historical data collected by one or more behavioural policies (experts) by interacting with an environment. However, the individual experts may be privacy-sensitive in that the learnt policy may retain information about their precise choices. In some domains like personalized retrieval, advertising and healthcare, the expert choices are considered sensitive data. To provably protect the privacy of such experts, we propose a novel consensus-based expert-level differentially private offline RL training approach compatible with any existing offline RL algorithm. We prove rigorous differential privacy guarantees, while maintaining strong empirical performance. Unlike existing work in differentially private RL, we supplement the theory with proof-of-concept experiments on classic RL environments featuring large continuous state spaces, demonstrating substantial improvements over a natural baseline across multiple tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13598
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Preserving Expert-Level Privacy in Offline Reinforcement Learning
Sharma, Navodita
Vinod, Vishnu
Thakurta, Abhradeep
Agarwal, Alekh
Balle, Borja
Dann, Christoph
Raghuveer, Aravindan
Cryptography and Security
Machine Learning
The offline reinforcement learning (RL) problem aims to learn an optimal policy from historical data collected by one or more behavioural policies (experts) by interacting with an environment. However, the individual experts may be privacy-sensitive in that the learnt policy may retain information about their precise choices. In some domains like personalized retrieval, advertising and healthcare, the expert choices are considered sensitive data. To provably protect the privacy of such experts, we propose a novel consensus-based expert-level differentially private offline RL training approach compatible with any existing offline RL algorithm. We prove rigorous differential privacy guarantees, while maintaining strong empirical performance. Unlike existing work in differentially private RL, we supplement the theory with proof-of-concept experiments on classic RL environments featuring large continuous state spaces, demonstrating substantial improvements over a natural baseline across multiple tasks.
title Preserving Expert-Level Privacy in Offline Reinforcement Learning
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2411.13598