On the continuity and smoothness of the value function in reinforcement learning and optimal control
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929284720885760 |
|---|---|
| author | Harder, Hans Peitz, Sebastian |
| author_facet | Harder, Hans Peitz, Sebastian |
| contents | The value function plays a crucial role as a measure for the cumulative future reward an agent receives in both reinforcement learning and optimal control. It is therefore of interest to study how similar the values of neighboring states are, i.e., to investigate the continuity of the value function. We do so by providing and verifying upper bounds on the value function's modulus of continuity. Additionally, we show that the value function is always Hölder continuous under relatively weak assumptions on the underlying system and that non-differentiable value functions can be made differentiable by slightly "disturbing" the system. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_14432 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | On the continuity and smoothness of the value function in reinforcement learning and optimal control Harder, Hans Peitz, Sebastian Systems and Control Artificial Intelligence 37H99, 37N35, 93E03 I.2.8 The value function plays a crucial role as a measure for the cumulative future reward an agent receives in both reinforcement learning and optimal control. It is therefore of interest to study how similar the values of neighboring states are, i.e., to investigate the continuity of the value function. We do so by providing and verifying upper bounds on the value function's modulus of continuity. Additionally, we show that the value function is always Hölder continuous under relatively weak assumptions on the underlying system and that non-differentiable value functions can be made differentiable by slightly "disturbing" the system. |
| title | On the continuity and smoothness of the value function in reinforcement learning and optimal control |
| topic | Systems and Control Artificial Intelligence 37H99, 37N35, 93E03 I.2.8 |
| url | https://arxiv.org/abs/2403.14432 |