Entropy-regularized Point-based Value Iteration
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929243420622848 |
|---|---|
| author | Delecki, Harrison Vazquez-Chanlatte, Marcell Yel, Esen Wray, Kyle Arnon, Tomer Witwicki, Stefan Kochenderfer, Mykel J. |
| author_facet | Delecki, Harrison Vazquez-Chanlatte, Marcell Yel, Esen Wray, Kyle Arnon, Tomer Witwicki, Stefan Kochenderfer, Mykel J. |
| contents | Model-based planners for partially observable problems must accommodate both model uncertainty during planning and goal uncertainty during objective inference. However, model-based planners may be brittle under these types of uncertainty because they rely on an exact model and tend to commit to a single optimal behavior. Inspired by results in the model-free setting, we propose an entropy-regularized model-based planner for partially observable problems. Entropy regularization promotes policy robustness for planning and objective inference by encouraging policies to be no more committed to a single action than necessary. We evaluate the robustness and objective inference performance of entropy-regularized policies in three problem domains. Our results show that entropy-regularized policies outperform non-entropy-regularized baselines in terms of higher expected returns under modeling errors and higher accuracy during objective inference. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_09388 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Entropy-regularized Point-based Value Iteration Delecki, Harrison Vazquez-Chanlatte, Marcell Yel, Esen Wray, Kyle Arnon, Tomer Witwicki, Stefan Kochenderfer, Mykel J. Artificial Intelligence Model-based planners for partially observable problems must accommodate both model uncertainty during planning and goal uncertainty during objective inference. However, model-based planners may be brittle under these types of uncertainty because they rely on an exact model and tend to commit to a single optimal behavior. Inspired by results in the model-free setting, we propose an entropy-regularized model-based planner for partially observable problems. Entropy regularization promotes policy robustness for planning and objective inference by encouraging policies to be no more committed to a single action than necessary. We evaluate the robustness and objective inference performance of entropy-regularized policies in three problem domains. Our results show that entropy-regularized policies outperform non-entropy-regularized baselines in terms of higher expected returns under modeling errors and higher accuracy during objective inference. |
| title | Entropy-regularized Point-based Value Iteration |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2402.09388 |