Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915343477243904 |
|---|---|
| author | Suh, Jihoon Jang, Yeongjun Teranishi, Kaoru Tanaka, Takashi |
| author_facet | Suh, Jihoon Jang, Yeongjun Teranishi, Kaoru Tanaka, Takashi |
| contents | We propose an efficient encrypted policy synthesis to develop privacy-preserving model-based reinforcement learning. We first demonstrate that the relative-entropy-regularized reinforcement learning framework offers a computationally convenient linear and ``min-free'' structure for value iteration, enabling a direct and efficient integration of fully homomorphic encryption with bootstrapping into policy synthesis. Convergence and error bounds are analyzed as encrypted policy synthesis propagates errors under the presence of encryption-induced errors including quantization and bootstrapping. Theoretical analysis is validated by numerical simulations. Results demonstrate the effectiveness of the RERL framework in integrating FHE for encrypted policy synthesis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_12358 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis Suh, Jihoon Jang, Yeongjun Teranishi, Kaoru Tanaka, Takashi Machine Learning Systems and Control We propose an efficient encrypted policy synthesis to develop privacy-preserving model-based reinforcement learning. We first demonstrate that the relative-entropy-regularized reinforcement learning framework offers a computationally convenient linear and ``min-free'' structure for value iteration, enabling a direct and efficient integration of fully homomorphic encryption with bootstrapping into policy synthesis. Convergence and error bounds are analyzed as encrypted policy synthesis propagates errors under the presence of encryption-induced errors including quantization and bootstrapping. Theoretical analysis is validated by numerical simulations. Results demonstrate the effectiveness of the RERL framework in integrating FHE for encrypted policy synthesis. |
| title | Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis |
| topic | Machine Learning Systems and Control |
| url | https://arxiv.org/abs/2506.12358 |