Differentially Private Policy Gradient

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rio, Alexandre, Barlier, Merwan, Colin, Igor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910807254630400
author Rio, Alexandre
Barlier, Merwan
Colin, Igor
author_facet Rio, Alexandre
Barlier, Merwan
Colin, Igor
contents Motivated by the increasing deployment of reinforcement learning in the real world, involving a large consumption of personal data, we introduce a differentially private (DP) policy gradient algorithm. We show that, in this setting, the introduction of Differential Privacy can be reduced to the computation of appropriate trust regions, thus avoiding the sacrifice of theoretical properties of the DP-less methods. Therefore, we show that it is possible to find the right trade-off between privacy noise and trust-region size to obtain a performant differentially private policy gradient algorithm. We then outline its performance empirically on various benchmarks. Our results and the complexity of the tasks addressed represent a significant improvement over existing DP algorithms in online RL.
format Preprint
id arxiv_https___arxiv_org_abs_2501_19080
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Differentially Private Policy Gradient
Rio, Alexandre
Barlier, Merwan
Colin, Igor
Machine Learning
Motivated by the increasing deployment of reinforcement learning in the real world, involving a large consumption of personal data, we introduce a differentially private (DP) policy gradient algorithm. We show that, in this setting, the introduction of Differential Privacy can be reduced to the computation of appropriate trust regions, thus avoiding the sacrifice of theoretical properties of the DP-less methods. Therefore, we show that it is possible to find the right trade-off between privacy noise and trust-region size to obtain a performant differentially private policy gradient algorithm. We then outline its performance empirically on various benchmarks. Our results and the complexity of the tasks addressed represent a significant improvement over existing DP algorithms in online RL.
title Differentially Private Policy Gradient
topic Machine Learning
url https://arxiv.org/abs/2501.19080