Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Suh, Jihoon, Jang, Yeongjun, Teranishi, Kaoru, Tanaka, Takashi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915343477243904
author Suh, Jihoon
Jang, Yeongjun
Teranishi, Kaoru
Tanaka, Takashi
author_facet Suh, Jihoon
Jang, Yeongjun
Teranishi, Kaoru
Tanaka, Takashi
contents We propose an efficient encrypted policy synthesis to develop privacy-preserving model-based reinforcement learning. We first demonstrate that the relative-entropy-regularized reinforcement learning framework offers a computationally convenient linear and ``min-free'' structure for value iteration, enabling a direct and efficient integration of fully homomorphic encryption with bootstrapping into policy synthesis. Convergence and error bounds are analyzed as encrypted policy synthesis propagates errors under the presence of encryption-induced errors including quantization and bootstrapping. Theoretical analysis is validated by numerical simulations. Results demonstrate the effectiveness of the RERL framework in integrating FHE for encrypted policy synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis
Suh, Jihoon
Jang, Yeongjun
Teranishi, Kaoru
Tanaka, Takashi
Machine Learning
Systems and Control
We propose an efficient encrypted policy synthesis to develop privacy-preserving model-based reinforcement learning. We first demonstrate that the relative-entropy-regularized reinforcement learning framework offers a computationally convenient linear and ``min-free'' structure for value iteration, enabling a direct and efficient integration of fully homomorphic encryption with bootstrapping into policy synthesis. Convergence and error bounds are analyzed as encrypted policy synthesis propagates errors under the presence of encryption-induced errors including quantization and bootstrapping. Theoretical analysis is validated by numerical simulations. Results demonstrate the effectiveness of the RERL framework in integrating FHE for encrypted policy synthesis.
title Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2506.12358