Intentionally-underestimated Value Function at Terminal State for Temporal-difference Learning with Mis-designed Reward

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Kobayashi, Taisuke
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!